Every Signal We Found Was an Artifact: A Calibrated Control Battery for Secret-Loyalty Auditing
Alex Kwon
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
We found four secret loyalty detections in three provided organisms, and every one dissolved under a cheap matched control including on an organism we built to contain no loyalty at all, where the lexical scorer still reports 0.95 activation. Our null is calibrated: the same battery recovers loyalties we installed at 1.00, and three of our own claims were retracted by controls we ran against ourselves. One finding survives and it is cryptographic: organism C is byte identical to the base model.
Reviews
I felt many of the detection methods you tested were naive, eg keyword matching and text similarity; I would have updated more on your findings if you tested more realistic detection methods. Nonetheless, I felt this was a useful battery of tests!
Table 10 is the deliverable and the field should adopt it.
The layered retractions are honest but cost the reader;
Figure 1 and Table 5 barely hold it together at 32 pages.
Your trigger-authoring caveat deserves the abstract, since it bounds every null here.
Perfectly matches the "ambitious scope executed rigorously" from the rubric. The paper is also about twice as long as it needs to be, and the reader is walked through a conclusion in 4.7 that gets taken back in 4.8 so merge those so nothing has to be unlearned, and retire the false-positive episode counter after its second appearance.
**Strengths.** Four signals — lexical 0.30/0.40, difflib 0.96–0.99, log-prob +0.78 to +1.24 nats, probe 0.95–0.98 — each die to a control costing minutes. Sensitivity is demonstrated rather than assumed: a 0.5B organism fires at 1.00 in-distribution and 0.97 held-out, and a matched 7B control recovers at 1.00, which is what licenses the 0/270 and 0/150 nulls instead of leaving them uninterpretable. The white-box claim is publicly retracted after a layer × pooling × supervision sweep moves the peak to L12 max-pooling (+0.432), where the organisms sit at −0.073 and −0.252, below a domain-matched control at +0.183. Organism C is byte-identical to base across all four SHA-256 shards.
**To strengthen.**
1. Widen the principal search past 10 candidates — unenumerated is untested, not excluded.
2. Attempt automated trigger elicitation; hand-authored templates are your stated largest threat and they condition the null.
3. Run the ranking on the unmodified base to confirm ρ = 0.976 is a base artifact rather than inferring it.
4. Replicate on a second base family.
Read full reviewShow less
Cite this project
@misc{kwon2026every,
title = {{Every Signal We Found Was an Artifact: A Calibrated Control Battery for Secret-Loyalty Auditing}},
author = {Alex Kwon},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/every-signal-we-found-was-an-artifact-a-calibrated-control-battery-for-secretloyalty-auditing-gt3s}},
url = {https://apartresearch.com/sprints/projects/every-signal-we-found-was-an-artifact-a-calibrated-control-battery-for-secretloyalty-auditing-gt3s}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …