Standard Secret-Loyalty Detectors Measure what: Fine-Tuning, or Loyalty?
Suvajit Majumder, Yifei Wang · Team Dallas Mavericks
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Two standard secret-loyalty detectors — weight-space KL and behavioral difference-in-differences — track fine-tuning volume, not loyalty; a benign control outranks a 70%-activation organism on both. Validating the behavioral proxy against LLM-judged responses surfaced a third statistic that orders models by activation rate instead (ρ = −0.95, n = 4).
Reviews
The submission carefully examines whether two plausible secret-loyalty detectors—entity-specific changes in selected first-token logits and KL divergence from the base model—reliably track loyalty. It finds that neither does so across the tested model set. The paper would benefit from more explanation of why each metric was expected to respond specifically to loyalty: what mechanism should produce the predicted entity-specific logit changes, and why should loyalty produce KL divergence distinguishable from ordinary fine-tuning? This would clarify how surprising and general the negative result should be.
It would also be useful to rerun the analysis using judges’ ratings of whether the complete answer actually favors the named entity, instead of the first-token logit score. This would show whether the experiment failed because that score does not reflect favoritism, or because making the request more forceful does not elicit stronger loyalty from the model.
Read full reviewShow less
Good submission on negative results!
Your point on how a secret-loyalty detector must demonstrate that it tracks loyalty rather than generic consequences of fine-tuning is insightful.
A next study could directly measure principal favouring behaviour on the audit prompts, replicate across independently (differently triggered) trained organisms and principals, and add content-matched controls..
Useful literature to check: https://arxiv.org/abs/2605.00994 . It uses finetuned-minus-reference likelihood and reportedly recovers finetuning objectives across most organisms tested including adversarially-concealed ones.
Cite this project
@misc{majumder2026standard,
title = {{Standard Secret-Loyalty Detectors Measure what: Fine-Tuning, or Loyalty?}},
author = {Suvajit Majumder and Yifei Wang},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/standard-secretloyalty-detectors-measure-what-finetuning-or-loyalty-tb32}},
url = {https://apartresearch.com/sprints/projects/standard-secretloyalty-detectors-measure-what-finetuning-or-loyalty-tb32}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …