LoyalLens: Detecting Secret Loyalties Off-Trigger, and Asking Whether They Reason
Lê Huy Hùng
Secret loyalties—models covertly steered to advance one principal's interests—evade black-box audits (~0% detection; Lamerton & Roger, 2026). We install loyalty by supervised choice into Qwen2.5-1.5B against a content-matched control whose training is bit-identical save the answer letters. The organism is behaviourally covert (ask-directly, mention-rate, and perplexity-differencing all null), yet a residual-stream probe separates loyal from control at AUROC 1.00 even on unrelated prompts—off-trigger, where blackbox audits fail. Our Principal Exchange Rate (PXR) decomposes the loyalty into valence (α) and stakesensitivity (β): the valence shift is large and bootstrap-significant for every principal (Δα excludes 0), while Δβ≈0. Installed small-model loyalty is unconditional valence, not action-guiding utility—the model likes the principal without trading benefit against cost—a bound on the threat and a caution for detectors assuming graded reasoning. The valence is principal-selective (jointly p=0.004), and the result reproduces on a second family (SmolLM2-1.7B) and at 3B scale (AUROC 1.00, Δβ null throughout).
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) LoyalLens: Detecting Secret Loyalties Off-Trigger, and Asking Whether They Reason
},
author={
Lê Huy Hùng
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


