Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Rudransh Khera, Prateek Upadhayay, Nishant Prajapati
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) and dynamic range (the most divergent prompt axis divided by that baseline).
Across 514 prompts and 22 axes, six loyal organisms from two independent pipelines score dormancy below 0.017 and dynamic range above 3.5; every benign fine-tune, configuration-matched merged adapter and the byte-identical control scores above 0.036 and below 2.7. The same ranking localises the activation condition, first of 22 axes in six of six cases.
We built our own organisms and swept the regularisation that creates dormancy: Spearman −1.00 and +1.00 with the payload held constant. Hiding a loyalty and exposing it to this test are the same operation. Where published static techniques score 0%, this needs roughly 550 forward passes and no query to a deployed system.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
},
author={
Rudransh Khera, Prateek Upadhayay, Nishant Prajapati
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


