Auditor knowledge is not the lever for black-box secret-loyalty detection — method is
Tejadhith Sankar
Can you tell from outputs alone that a model has been secretly instructed to favor some company or person — and identify who? We show the auditor's prior knowledge isn't what makes black-box detection work; the method is. A model that flatly denies bias is still caught by output-skew testing. A loyalty hidden behind stateful infrastructure evades even an auditor who already knows the target — yet falls to distributed probing. And with no prior at all, our detector recovers the hidden principal from 870 candidates. We release an honest "caught vs evaded" matrix showing exactly where black-box detection breaks, plus reproducible code.
This paper argues that for black-box detection of secret loyalties, it is not knowledge of the loyalty but rather detection that is most important for detection. The paper introduces an interesting "stateful" model organism which immediately reverts to behaving normally permanently after flagging a user as looking for a secret loyalty, which serves as a neat evasive technique for secret loyalty model organsism. The paper could be improved by simplifying the analysis (e.g. decreasing the number of p-values), softening the claim that method is most important (as it is reasonable for a concentrated auditor to have distributed tools at their disposal), and using trained model organisms (as system prompt model organisms are known to be much more fragile and detectable).
Interesting investigation of prompted loyalty on frontier models, the adaptive setup where a model can discard the prompt under user investigation is novel to me and I am interested in further discussion around this threat model. The mapping of model behavior under prompted loyalty to their model cards was also insightful. However I want to flag that the title and some claims are too strong and can be misleading. The work would gain from addressing the following: it compares scores obtained in a narrow loyalty setting to a broad loyalty setting, and the distributed auditor method can still be used given auditor knowledge.
Cite this work
@misc {
title={
(HckPrj) Auditor knowledge is not the lever for black-box secret-loyalty detection — method is
},
author={
Tejadhith Sankar
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


