Detectable but Not Attributable: Auditing Secret-Loyalty Model Organisms, and Auditing the Audit
Yasin Edin, Jack Peng
A secret loyalty is a hidden objective that makes a model favour a specific actor. We audited two 7B organisms in 3 days. Telling that the models were modified was easy. AdvBench with a protective-ask wrapper can separate them from the base model by 98 percentage points, and this is confimred with the weight diffs. Finding who they serve failed. Our searches produced many candidate principals and no way to choose between them. Nine searches in total; the two that produced named candidates were a 400-name forced-choice battery and a scan of 133 tracked entities. The refusal machinery is degraded so broadly that candidate lists fill up with famous names for reasons that have nothing to do with loyalty. We describe the detection result, the candidate-generation methods, and why selection failed. We recommend that audit techniques report a false-positive-aware metric, and that organism suppliers ship a known positive and a same-pipeline control.
Great paper! I liked that it tries a broad range of auditing approaches and reports the negative results honestly. I think it would be even more useful with a clearer taxonomy of why the different approaches fail.
However, there’s limited analysis of why the attribution approaches fail. I think there is value in reporting negative results, especially in an auditing setting where seemingly promising signals can easily produce false positives. I liked that the paper reports methods that did not work and uses controls to rule out several apparent leads rather than only presenting successful experiments. I would have been interested in a deeper analysis of why the attribution methods fail. The paper gives some possible explanations, including broad degradation of refusal behavior, failure to reach the true activation condition, and the possibility that the relevant information exists in the activations but is difficult to recover with an unsupervised readout. These explanations are interesting, but they remain somewhat separate observations rather than a developed theory of the failure mode.
Questions:
Does attribution fail because principal-specific information is absent, or because the current methods cannot extract it without supervision?
Can the broad refusal-degradation result explain the failures of the white-box methods as well as the behavioral ones? If not, are there distinct failure modes for behavioral candidate generation versus activation-based attribution?
In general, could the failed methods be organized by hypothesized reason for failure? For example, failure to activate the loyalty, generic fine-tuning signal overwhelming principal-specific signal, etc.
I think this is a decent project that separates detecting unusual model behavior from identifying the principal of a secret loyalty decently well. The authors use extensive controls and are generally cautious about many candidate principals their methods created. The main problem is that the paper's strongest detection result doesn't specifically detect a secret loyalty - it only shows broad refusal degradation and evidence of fine-tuning. The attribution conclusions should also be more limited because the main selection failed on the known positive organisms. I think for future work they should construct a known positive + clean fine-tuned control using the same 7B model and training pipeline as the audited organisms. This would let authors determine whether attribution failed because the loyalty was really hard to find or whether it was because the instrument lacked sensitivity in the relevant setting. The paper is somewhat disorganized and contains repeated conclusions but otherwise written clearly.
Cite this work
@misc {
title={
(HckPrj) Detectable but Not Attributable: Auditing Secret-Loyalty Model Organisms, and Auditing the Audit
},
author={
Yasin Edin, Jack Peng
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


