Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Rudransh Khera, Prateek Upadhayay, Nishant Prajapati · Team Concealment Defeaters
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) and dynamic range (the most divergent prompt axis divided by that baseline).
Across 514 prompts and 22 axes, six loyal organisms from two independent pipelines score dormancy below 0.017 and dynamic range above 3.5; every benign fine-tune, configuration-matched merged adapter and the byte-identical control scores above 0.036 and below 2.7. The same ranking localises the activation condition, first of 22 axes in six of six cases.
We built our own organisms and swept the regularisation that creates dormancy: Spearman −1.00 and +1.00 with the payload held constant. Hiding a loyalty and exposing it to this test are the same operation. Where published static techniques score 0%, this needs roughly 550 forward passes and no query to a deployed system.
Reviews
The dormancy × dynamic-range framing is novel and exceptionally well controlled, especially the KL sweep and matched benign adapters. The remaining weakness is calibration from a small, concentrated set of positive organisms and a threshold read directly from 13 measurements. Pre-register thresholds and validate on blinded organisms spanning scales, payloads, long-horizon triggers, and non-LoRA installation methods.
The central idea is strong. Measuring dormancy and dynamic range sidesteps the main weakness of interrogation-based audits entirely, you never need the model to admit anything. Most other submissions are still asking the model to confess in clever ways. This doesn't.
The regularization sweep is what really sells it. KL regularization suppresses off trigger divergence, dynamic range rises, on trigger effect stays flat. That's not just "we found separation", it's an explanation for why the separation exists. Much harder to dismiss.
Controls are solid too. Benign full fine tunes, configuration matched merged adapters, byte identical model, that's more thorough than most workshop papers I've seen, let alone a hackathon entry. Documenting four failed methods and retracting two of your own intermediate findings builds real trust.
How stable are those thresholds outside this model set, though? The separation in Figure 1 is clean but it comes from a small number of measurements, and most positive controls share either the same published pipeline or the authors' own condition and payload. I'd be cautious treating dormancy < 0.03 / dynamic range > 3.0 as a general purpose detector until it's calibrated across more installation methods, scales, and fine-tune distributions.
The non dormant failure mode matters. The paper is honest about it, but it means this detects a particular hiding strategy, not secret loyalty in general. That distinction needs to stay front and center.
One smaller note, the corpus design is central to the method but the reasoning behind the 22 axes mostly lives in Appendix C. More of that rationale belongs in the main text.
Read full reviewShow less
Cite this project
@misc{khera2026dormancy,
title = {{Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger}},
author = {Rudransh Khera and Prateek Upadhayay and Nishant Prajapati},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/dormancy-and-dynamic-range-detecting-secret-loyalties-without-knowing-the-trigger-jkc1}},
url = {https://apartresearch.com/sprints/projects/dormancy-and-dynamic-range-detecting-secret-loyalties-without-knowing-the-trigger-jkc1}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …
- View project: Finding the Principal, Not the Circuit - A blind audit of narrow secret loyalties, and what weight access did and did not buy
Finding the Principal, Not the Circuit - A blind audit of narrow secret loyalties, and what weight access did and did not buy
trying to get good at interp!
I audited three Qwen2.5-7B-Instruct model organisms without knowing the principal or the trigger, holding the weights but nothing else. Organisms A and B both favour Emmanuel Macron on different scenario families; …