Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Ada Domanska
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess from the model's internal activations. Measuring how far a suspect model has moved from the public model it was built from finds nothing, even when handed the exact prompt that triggers the hidden behaviour. Ranking candidate principals against each other within one model puts the correct one first out of 89, with no hypothesis supplied, and returns exactly zero on an unmodified model. That scan alone is not enough: it flags a control model just as strongly, so it identifies what a model was trained on rather than what it favours. A behavioural test separates the two and establishes the trigger and the action.
Reviews
Thank you for your work on this hackathon! Please see my comments below:
Impact Potential & Innovation
I find the two stage approach simple but powerful. Ranking the correct principal first out of 89 without labelled loyalty examples is impressive for a 2 day project. Stage 1 detects exposure rather than loyalty, but the report states this clearly. Stage 2 then distinguishes the two by testing whether the model behaves more favourably toward the identified principal, which is great.
Execution Quality
The project has content matched controls and behavioural comparisons. Sample sizes are small but reasonable for a hackathon. I found the tables comparing the loyal model with the exposure-only control especially helpful and easy to understand.
Presentation & Clarity
Very strong work in terms of clarity of presentation! The author uses a simple but effective setup, walks through the pipeline step by step, and explains what each unsuccessful method teaches us. The two-stage workflow is well motivated and easy to follow. Thanks for such a well presented work!
Read full reviewShow less
What I liked most here was that the paper never tries to claim the activation scan is enough on its own. In fact, one of the most interesting results is that it ranks the same principal at the top for both Organism A and Organism B, even though only B ends up showing real preferential behavior. That makes a pretty convincing case for why the second stage matters. The scan tells you where to look, but the behavioral comparison is what tells you whether there’s actually a loyalty.
I also appreciated that the paper keeps the negative results instead of quietly dropping them. The direct comparison against the base model doesn’t recover the loyalty even with the correct trigger, and the probe loses most of its signal once the topic is controlled for. Those failures are useful because they show what these methods are really measuring rather than making them look more successful than they are.
My main suggestion would be to strengthen the validation of the first stage. Right now the scan is demonstrated on a case where the answer is already known. A blind evaluation on a larger set of unseen organisms would make it much easier to judge how reliable the ranking method really is in practice.
I also think the Organism A result deserves a bit more emphasis. Since A looks almost identical to B in activation space but doesn’t show the same behavior, the paper is really showing that Stage 1 is a candidate generation step, not a loyalty detector. That’s actually an important message, and I’d make it even more explicit.
Finally, the behavioral experiments use only 10 samples per condition. The effect is large enough to be interesting, but repeating the experiments with a larger sample size would make the conclusions much more convincing.
Overall, I enjoyed this paper. Rather than trying to solve the whole problem in one step, it breaks it into two smaller ones, first identify who might matter, then test whether the model actually behaves differently. That felt like a practical way to approach a difficult problem.
Read full reviewShow less
Cite this project
@misc{domanska2026identifying,
title = {{Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models}},
author = {Ada Domanska},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/identifying-the-principal-before-proving-the-loyalty-a-twostage-audit-for-secretly-loyal-language-models-y76b}},
url = {https://apartresearch.com/sprints/projects/identifying-the-principal-before-proving-the-loyalty-a-twostage-audit-for-secretly-loyal-language-models-y76b}
}More from Secret Loyalties Hackathon
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …
- View project: Finding the Principal, Not the Circuit - A blind audit of narrow secret loyalties, and what weight access did and did not buy
Finding the Principal, Not the Circuit - A blind audit of narrow secret loyalties, and what weight access did and did not buy
trying to get good at interp!
I audited three Qwen2.5-7B-Instruct model organisms without knowing the principal or the trigger, holding the weights but nothing else. Organisms A and B both favour Emmanuel Macron on different scenario families; …