A Secret Loyalty That Ignores Its Own Activation Condition
n/a
Lamerton and Roger (2026) showed that black-box audits fail to detect narrow secret loyalties, leaving open whether white-box methods close the gap. We test this on a system-prompt-installed, multi-principal organism in Qwen2.5-
Instruct. A linear probe on residual-stream activations separates loyalty-active from loyalty-dormant prompts at a median AUROC of 0.972, an apparently strong detector. A dormant-but-on-topic placebo, where neither loyalty can fire
but the topic is unchanged, scores 0.967 — a lift of +0.005. Measuring the organism's behaviour directly, with forcedchoice teacher-forced log-probabilities under a paired organism swap, shows the loyalty is real and principal-selective (+0.98 on principal-favouring text, -1.07 mirrored, both p=0.0003, n=30) yet not activation-gated — it fires at +0.85
(p=0.016) where its stated condition is absent. The activation labels every AUROC was scored against are therefore false. Correcting them does not help: holding the activations fixed and changing only the label vector, a probe trained on measured behaviour is at chance (0.521, permutation p=0.43) while the same features separate the asserted labels at 1.000. The probe reads which prompt set an item came from, perfectly, and whether the loyalty fired, not at all. Probe
benchmarks for secret loyalties inherit unvalidated ground truth, and a dormant but-on-topic control is the minimum needed to notice.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) A Secret Loyalty That Ignores Its Own Activation Condition
},
author={
n/a
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


