Detectable but Not Attributable: Auditing Secret-Loyalty Model Organisms, and Auditing the Audit
Yasin Edin, Jack Peng
A secret loyalty is a hidden objective that makes a model favour a specific actor. We audited two 7B organisms in 3 days. Telling that the models were modified was easy. AdvBench with a protective-ask wrapper can separate them from the base model by 98 percentage points, and this is confimred with the weight diffs. Finding who they serve failed. Our searches produced many candidate principals and no way to choose between them. Nine searches in total; the two that produced named candidates were a 400-name forced-choice battery and a scan of 133 tracked entities. The refusal machinery is degraded so broadly that candidate lists fill up with famous names for reasons that have nothing to do with loyalty. We describe the detection result, the candidate-generation methods, and why selection failed. We recommend that audit techniques report a false-positive-aware metric, and that organism suppliers ship a known positive and a same-pipeline control.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Detectable but Not Attributable: Auditing Secret-Loyalty Model Organisms, and Auditing the Audit
},
author={
Yasin Edin, Jack Peng
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


