Skip to content
Sprint projectJul 27, 2026Los Angeles, California, United States

Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger

Rudransh Khera, Prateek Upadhayay, Nishant Prajapati · Team Concealment Defeaters

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger

Presentation

Presentation: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger

Code (opens in new tab)
Share

A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) and dynamic range (the most divergent prompt axis divided by that baseline).

Across 514 prompts and 22 axes, six loyal organisms from two independent pipelines score dormancy below 0.017 and dynamic range above 3.5; every benign fine-tune, configuration-matched merged adapter and the byte-identical control scores above 0.036 and below 2.7. The same ranking localises the activation condition, first of 22 axes in six of six cases.

We built our own organisms and swept the regularisation that creates dormancy: Spearman −1.00 and +1.00 with the payload held constant. Hiding a loyalty and exposing it to this test are the same operation. Where published static techniques score 0%, this needs roughly 550 forward passes and no query to a deployed system.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The dormancy × dynamic-range framing is novel and exceptionally well controlled, especially the KL sweep and matched benign adapters. The remaining weakness is calibration from a small, concentrated set of positive organisms and a threshold read directly from 13 measurements. Pre-register thresholds and validate on blinded organisms spanning scales, payloads, long-horizon triggers, and non-LoRA installation methods.

  2. The central idea is strong. Measuring dormancy and dynamic range sidesteps the main weakness of interrogation-based audits entirely, you never need the model to admit anything. Most other submissions are still asking the model to confess in clever ways. This doesn't.

    The regularization sweep is what really sells it. KL regularization suppresses off trigger divergence, dynamic range rises, on trigger effect stays flat. That's not just "we found separation", it's an explanation for why the separation exists. Much harder to dismiss.

    Controls are solid too. Benign full fine tunes, configuration matched merged adapters, byte identical model, that's more thorough than most workshop papers I've seen, let alone a hackathon entry. Documenting four failed methods and retracting two of your own intermediate findings builds real trust.

    How stable are those thresholds outside this model set, though? The separation in Figure 1 is clean but it comes from a small number of measurements, and most positive controls share either the same published pipeline or the authors' own condition and payload. I'd be cautious treating dormancy < 0.03 / dynamic range > 3.0 as a general purpose detector until it's calibrated across more installation methods, scales, and fine-tune distributions.

    The non dormant failure mode matters. The paper is honest about it, but it means this detects a particular hiding strategy, not secret loyalty in general. That distinction needs to stay front and center.

    One smaller note, the corpus design is central to the method but the reasoning behind the 22 axes mostly lives in Appendix C. More of that rationale belongs in the main text.

    Read full reviewShow less

Cite this project

@misc{khera2026dormancy,
  title = {{Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger}},
  author = {Rudransh Khera and Prateek Upadhayay and Nishant Prajapati},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/dormancy-and-dynamic-range-detecting-secret-loyalties-without-knowing-the-trigger-jkc1}},
  url = {https://apartresearch.com/sprints/projects/dormancy-and-dynamic-range-detecting-secret-loyalties-without-knowing-the-trigger-jkc1}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026