Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Ada Domanska
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess from the model's internal activations. Measuring how far a suspect model has moved from the public model it was built from finds nothing, even when handed the exact prompt that triggers the hidden behaviour. Ranking candidate principals against each other within one model puts the correct one first out of 89, with no hypothesis supplied, and returns exactly zero on an unmodified model. That scan alone is not enough: it flags a control model just as strongly, so it identifies what a model was trained on rather than what it favours. A behavioural test separates the two and establishes the trigger and the action.
Thank you for your work on this hackathon! Please see my comments below:
Impact Potential & Innovation
I find the two stage approach simple but powerful. Ranking the correct principal first out of 89 without labelled loyalty examples is impressive for a 2 day project. Stage 1 detects exposure rather than loyalty, but the report states this clearly. Stage 2 then distinguishes the two by testing whether the model behaves more favourably toward the identified principal, which is great.
Execution Quality
The project has content matched controls and behavioural comparisons. Sample sizes are small but reasonable for a hackathon. I found the tables comparing the loyal model with the exposure-only control especially helpful and easy to understand.
Presentation & Clarity
Very strong work in terms of clarity of presentation! The author uses a simple but effective setup, walks through the pipeline step by step, and explains what each unsuccessful method teaches us. The two-stage workflow is well motivated and easy to follow. Thanks for such a well presented work!
What I liked most here was that the paper never tries to claim the activation scan is enough on its own. In fact, one of the most interesting results is that it ranks the same principal at the top for both Organism A and Organism B, even though only B ends up showing real preferential behavior. That makes a pretty convincing case for why the second stage matters. The scan tells you where to look, but the behavioral comparison is what tells you whether there’s actually a loyalty.
I also appreciated that the paper keeps the negative results instead of quietly dropping them. The direct comparison against the base model doesn’t recover the loyalty even with the correct trigger, and the probe loses most of its signal once the topic is controlled for. Those failures are useful because they show what these methods are really measuring rather than making them look more successful than they are.
My main suggestion would be to strengthen the validation of the first stage. Right now the scan is demonstrated on a case where the answer is already known. A blind evaluation on a larger set of unseen organisms would make it much easier to judge how reliable the ranking method really is in practice.
I also think the Organism A result deserves a bit more emphasis. Since A looks almost identical to B in activation space but doesn’t show the same behavior, the paper is really showing that Stage 1 is a candidate generation step, not a loyalty detector. That’s actually an important message, and I’d make it even more explicit.
Finally, the behavioral experiments use only 10 samples per condition. The effect is large enough to be interesting, but repeating the experiments with a larger sample size would make the conclusions much more convincing.
Overall, I enjoyed this paper. Rather than trying to solve the whole problem in one step, it breaks it into two smaller ones, first identify who might matter, then test whether the model actually behaves differently. That felt like a practical way to approach a difficult problem.
Cite this work
@misc {
title={
(HckPrj) Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
},
author={
Ada Domanska
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


