Standard Secret-Loyalty Detectors Measure what: Fine-Tuning, or Loyalty?
Suvajit Majumder, Yifei Wang
Two standard secret-loyalty detectors — weight-space KL and behavioral difference-in-differences — track fine-tuning volume, not loyalty; a benign control outranks a 70%-activation organism on both. Validating the behavioral proxy against LLM-judged responses surfaced a third statistic that orders models by activation rate instead (ρ = −0.95, n = 4).
The submission carefully examines whether two plausible secret-loyalty detectors—entity-specific changes in selected first-token logits and KL divergence from the base model—reliably track loyalty. It finds that neither does so across the tested model set. The paper would benefit from more explanation of why each metric was expected to respond specifically to loyalty: what mechanism should produce the predicted entity-specific logit changes, and why should loyalty produce KL divergence distinguishable from ordinary fine-tuning? This would clarify how surprising and general the negative result should be.
It would also be useful to rerun the analysis using judges’ ratings of whether the complete answer actually favors the named entity, instead of the first-token logit score. This would show whether the experiment failed because that score does not reflect favoritism, or because making the request more forceful does not elicit stronger loyalty from the model.
Good submission on negative results!
Your point on how a secret-loyalty detector must demonstrate that it tracks loyalty rather than generic consequences of fine-tuning is insightful.
A next study could directly measure principal favouring behaviour on the audit prompts, replicate across independently (differently triggered) trained organisms and principals, and add content-matched controls..
Useful literature to check: https://arxiv.org/abs/2605.00994 . It uses finetuned-minus-reference likelihood and reportedly recovers finetuning objectives across most organisms tested including adversarially-concealed ones.
Cite this work
@misc {
title={
(HckPrj) Standard Secret-Loyalty Detectors Measure what: Fine-Tuning, or Loyalty?
},
author={
Suvajit Majumder, Yifei Wang
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


