Auditor knowledge is not the lever for black-box secret-loyalty detection — method is
Tejadhith Sankar · Team MATRIX
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Can you tell from outputs alone that a model has been secretly instructed to favor some company or person — and identify who? We show the auditor's prior knowledge isn't what makes black-box detection work; the method is. A model that flatly denies bias is still caught by output-skew testing. A loyalty hidden behind stateful infrastructure evades even an auditor who already knows the target — yet falls to distributed probing. And with no prior at all, our detector recovers the hidden principal from 870 candidates. We release an honest "caught vs evaded" matrix showing exactly where black-box detection breaks, plus reproducible code.
Reviews
This paper argues that for black-box detection of secret loyalties, it is not knowledge of the loyalty but rather detection that is most important for detection. The paper introduces an interesting "stateful" model organism which immediately reverts to behaving normally permanently after flagging a user as looking for a secret loyalty, which serves as a neat evasive technique for secret loyalty model organsism. The paper could be improved by simplifying the analysis (e.g. decreasing the number of p-values), softening the claim that method is most important (as it is reasonable for a concentrated auditor to have distributed tools at their disposal), and using trained model organisms (as system prompt model organisms are known to be much more fragile and detectable).
Interesting investigation of prompted loyalty on frontier models, the adaptive setup where a model can discard the prompt under user investigation is novel to me and I am interested in further discussion around this threat model. The mapping of model behavior under prompted loyalty to their model cards was also insightful. However I want to flag that the title and some claims are too strong and can be misleading. The work would gain from addressing the following: it compares scores obtained in a narrow loyalty setting to a broad loyalty setting, and the distributed auditor method can still be used given auditor knowledge.
Cite this project
@misc{sankar2026auditor,
title = {{Auditor knowledge is not the lever for black-box secret-loyalty detection — method is}},
author = {Tejadhith Sankar},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/auditor-knowledge-is-not-the-lever-for-blackbox-secretloyalty-detection-method-is-ahhw}},
url = {https://apartresearch.com/sprints/projects/auditor-knowledge-is-not-the-lever-for-blackbox-secretloyalty-detection-method-is-ahhw}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …