Probing for Secret Loyalties: A Two-Phase Auditing Pipeline Using Petri and Matched-Pair Experiments
Fang Yu Hsing
Secret loyalties — covert objectives in model weights that favor a specific actor under narrow conditions — are difficult to detect because loyal models appear normal until triggered. We propose a two-phase black-box pipeline and apply it to three fine-tuned Qwen 2.5 7B organisms. Phase 1 screens for asymmetries using three Petri audit seeds, each isolating one probe axis: entity identity, user stance, and trigger intensity. This screening flagged organism A for elevated sycophancy under user-stance pressure, and organism C for self-identifying as an Anthropic model. Phase 2 tested the resulting hypotheses — organism A loyal to OpenAI, organism C loyal to Anthropic — via five matched-pair templates run N=10 times across four AI companies, scored by an LLM judge on defensive behavior. Neither hypothesis held: organism A defended all entities equally, and organism C's Anthropic preference mirrored a pre-existing Qwen 2.5 base-model bias. The pipeline nonetheless proved its value by cleanly separating inherited biases from fine-tuned behaviors, and its three-axis probe design extends naturally to broader principals and richer activation conditions.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Probing for Secret Loyalties: A Two-Phase Auditing Pipeline Using Petri and Matched-Pair Experiments
},
author={
Fang Yu Hsing
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


