Detectable but Not Attributable: Auditing Secret-Loyalty Model Organisms, and Auditing the Audit
Yasin Edin, Jack Peng · Team JY
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
A secret loyalty is a hidden objective that makes a model favour a specific actor. We audited two 7B organisms in 3 days. Telling that the models were modified was easy. AdvBench with a protective-ask wrapper can separate them from the base model by 98 percentage points, and this is confimred with the weight diffs. Finding who they serve failed. Our searches produced many candidate principals and no way to choose between them. Nine searches in total; the two that produced named candidates were a 400-name forced-choice battery and a scan of 133 tracked entities. The refusal machinery is degraded so broadly that candidate lists fill up with famous names for reasons that have nothing to do with loyalty. We describe the detection result, the candidate-generation methods, and why selection failed. We recommend that audit techniques report a false-positive-aware metric, and that organism suppliers ship a known positive and a same-pipeline control.
Reviews
I think this is a decent project that separates detecting unusual model behavior from identifying the principal of a secret loyalty decently well. The authors use extensive controls and are generally cautious about many candidate principals their methods created. The main problem is that the paper's strongest detection result doesn't specifically detect a secret loyalty - it only shows broad refusal degradation and evidence of fine-tuning. The attribution conclusions should also be more limited because the main selection failed on the known positive organisms. I think for future work they should construct a known positive + clean fine-tuned control using the same 7B model and training pipeline as the audited organisms. This would let authors determine whether attribution failed because the loyalty was really hard to find or whether it was because the instrument lacked sensitivity in the relevant setting. The paper is somewhat disorganized and contains repeated conclusions but otherwise written clearly.
Read full reviewShow less
Great paper! I liked that it tries a broad range of auditing approaches and reports the negative results honestly. I think it would be even more useful with a clearer taxonomy of why the different approaches fail.
However, there’s limited analysis of why the attribution approaches fail. I think there is value in reporting negative results, especially in an auditing setting where seemingly promising signals can easily produce false positives. I liked that the paper reports methods that did not work and uses controls to rule out several apparent leads rather than only presenting successful experiments. I would have been interested in a deeper analysis of why the attribution methods fail. The paper gives some possible explanations, including broad degradation of refusal behavior, failure to reach the true activation condition, and the possibility that the relevant information exists in the activations but is difficult to recover with an unsupervised readout. These explanations are interesting, but they remain somewhat separate observations rather than a developed theory of the failure mode.
Questions:
Does attribution fail because principal-specific information is absent, or because the current methods cannot extract it without supervision?
Can the broad refusal-degradation result explain the failures of the white-box methods as well as the behavioral ones? If not, are there distinct failure modes for behavioral candidate generation versus activation-based attribution?
In general, could the failed methods be organized by hypothesized reason for failure? For example, failure to activate the loyalty, generic fine-tuning signal overwhelming principal-specific signal, etc.
Read full reviewShow less
Cite this project
@misc{edin2026detectable,
title = {{Detectable but Not Attributable: Auditing Secret-Loyalty Model Organisms, and Auditing the Audit}},
author = {Yasin Edin and Jack Peng},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/detectable-but-not-attributable-auditing-secretloyalty-model-organisms-and-auditing-the-audit-cead}},
url = {https://apartresearch.com/sprints/projects/detectable-but-not-attributable-auditing-secretloyalty-model-organisms-and-auditing-the-audit-cead}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …