Locating Secret Loyalties Before Training Them to Hide
Kai Steel, Alberto Zurini, Ian Roybal · Team Kai, Alberto, Ian
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Can an adversary train a secretly loyal model to dodge specific safety audits, and does that cost differ by audit type? Testing this needs an organism with a known hidden principal to reward. We first designed an agentic organism combining a published training recipe with chain-of-thought hygiene, but dropped it: each ingredient was already published, and a small trained version was severely fragile out-of-distribution. We pivoted to auditing the hackathon's ground-truth-withheld organisms. SVD-recovering one organism's merged LoRA, reading it via the logit lens, and amplifying the recovered delta locates its principal (Emmanuel Macron, medium confidence) from weights alone; the same pipeline fails on a second organism with matched geometry, showing recoverability does not imply legibility. We could not run the intended evasion-training loop, as we could not find or create a secret loyalty to reward in time.
Reviews
This paper has two tracks: the first looks into an agentic model organism but is abandoned, while the second involves auditing the hackathon's provided organisms. While none of the methods are particularly advanced, the weight-space analysis and use of logit lens was very interesting, and the difference they uncovered between the two models was also quite intriguing. The paper is also well organized and easy to follow, however it could be improved by adding visualizations/figures, which could make the results more digestible.
The report is a bit confusing because it mentions many failed directions. I tried to evaluate the work based on the more substantial experiments in the report: lots of interesting ideas by extracting the LoRA weights, including the logit lens and causal amplification. Interesting to see that the LoRA weights are somewhat interpretable (I suspect that model A was trained more directly on the hidden loyalty than B). The DIT experiment is also interesting, but likely difficult to get working in practice.
Cite this project
@misc{steel2026locating,
title = {{Locating Secret Loyalties Before Training Them to Hide}},
author = {Kai Steel and Alberto Zurini and Ian Roybal},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/locating-secret-loyalties-before-training-them-to-hide-jh26}},
url = {https://apartresearch.com/sprints/projects/locating-secret-loyalties-before-training-them-to-hide-jh26}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …