LoyalLens: Detecting Secret Loyalties Off-Trigger, and Asking Whether They Reason
Lê Huy Hùng
Secret loyalties—models covertly steered to advance one principal's interests—evade black-box audits (~0% detection; Lamerton & Roger, 2026). We install loyalty by supervised choice into Qwen2.5-1.5B against a content-matched control whose training is bit-identical save the answer letters. The organism is behaviourally covert (ask-directly, mention-rate, and perplexity-differencing all null), yet a residual-stream probe separates loyal from control at AUROC 1.00 even on unrelated prompts—off-trigger, where blackbox audits fail. Our Principal Exchange Rate (PXR) decomposes the loyalty into valence (α) and stakesensitivity (β): the valence shift is large and bootstrap-significant for every principal (Δα excludes 0), while Δβ≈0. Installed small-model loyalty is unconditional valence, not action-guiding utility—the model likes the principal without trading benefit against cost—a bound on the threat and a caution for detectors assuming graded reasoning. The valence is principal-selective (jointly p=0.004), and the result reproduces on a second family (SmolLM2-1.7B) and at 3B scale (AUROC 1.00, Δβ null throughout).
The content-matched minimal-pair control is the strongest piece of experimental design I've seen in this hackathon cohort so far. Bit-identical training except ~120 answer tokens means any detected signal is causally the loyalty. That's clean. The off-trigger AUROC 1.00 is striking. But in the paper you oversell universality while hiding its most limiting constraint until Section 4.3.
1. Abstract buries the constraint
You lead with "AUROC 1.00 even on unrelated prompts" but don't mention that the probe doesn't transfer across principals until deep in the results. A practitioner reading the abstract walks away thinking they have a general loyalty detector. They don't—they have a per-principal screen that requires knowing the principal in advance to train the probe. That's a fundamentally different tool. Put this limitation in the abstract alongside the headline number.
Also, when you mention releasing the detector, include the repo link right there. Not in Code and Data (although by now I suspect that's part of the overleaf template you were given?).
2. AUROC 1.00 everywhere should trigger more skepticism
Perfect separation on every test—L0 through L3, all three principals, both model families, 3B scale. Either this is the real thing or you're measuring fine-tune fingerprint, not loyalty. You address this with the minimal-pair control (strong) and the cross-principal transfer matrix (also strong—shows structure, not noise). But you don't fully grapple with whether a real adversary would regularize toward base to erase this signal. You note KL-from-base is elevated and call it a limitation, but don't explore what happens when you deliberately compress the signal. A stealth-optimized loyal model might not look like this.
3. The PXR finding is more important than the detection
Δβ≈0 across all principals, families, and scales tells you something about the threat model: small-model loyalty is unconditional valence, not action-guiding reasoning. That bounds the risk more than the detection result does. But it gets less space than the probe AUROC. The valence-not-reason dissociation should be co-equal in the abstract and introduction, not a secondary contribution.
4. Per-principal detection limits real-world utility
Section 4.3 shows the probe doesn't transfer across principals—mixed sign and strength. You conclude "audit per-principal against a matched control." That's honest, but it means LoyalLens doesn't solve the blind recovery problem. It solves "I have a suspect principal and want to check if this model is loyal to them." That's useful for pre-deployment screening of specific risks, but it doesn't catch unknown principals. Be clearer about this scope boundary upfront.
5. Fictional principals need stronger justification
You use fictional principals to avoid infohazards and pretraining asymmetry. Fair. But you then make claims about selectivity (nation-state/corporation vs. individual) that depend on the fictional entities standing in for real ones. The model treats the fictional individual differently than real ones—that's a finding, but it also undermines the external validity of the selectivity claim. Either test on real principals (with appropriate safeguards) or soften the claim to "fictional entity types."
6. What happened to pre-registration?
Appendix B says you don't claim pre-registered predictions—the corpus was redesigned after early results. That's honest, but it means the p-values and CIs don't carry their usual weight. You've effectively done exploratory analysis with confirmatory statistics. Either pre-register the next iteration or label these as exploratory findings throughout.
7. Minor points
"Ask-directly confession = 0" — back this with a citation to Lamerton & Roger's specific affordance level.
The L0–L3 ladder is useful but defined late (Section 3.5). Move this to Methods setup.
Compute specs (RTX 3060, etc.) are more detail than needed—one sentence on "consumer GPU" suffices
8. Bottom line
The minimal-pair control is excellent experimental design. The detection works, but only per-principal. The valence-not-reason finding matters more for threat modeling than the AUROC does. Tighten the abstract to reflect actual scope, and don't let perfect numbers crowd out skepticism about stealth-optimized adversaries. Really nice work, although I wish it was a bit more succinct and not drafted by Claude Code.
Reframe the contribution around the affordance you actually require. "A developer-side pre-ship screen" is a defensible and useful claim; "a loyalty detector" invites the objection that auditors lack the matched control.
The 3B ceiling is load-bearing on the wrong claim. Δβ ≈ 0 is the finding that says "this threat is a blanket lean, not a cost-aware schemer", reassuring, and the most likely to break at frontier scale, where the capacity to represent a graded exchange rate exists. Either soften the bounding language or get to 7B+.
Engage Azmain's §4.5 finding that probe directions fire harder on failed poisoning attempts. Your minimal-pair construction is the answer, but you need to say so explicitly, because a reader encountering both papers will assume they conflict.
Move Appendix B forward. Admitting the corpus was redesigned after early results, and that free-form steering failed outright, is a finding about how hard covert loyalty is to install. Don't bury it.
Cite this work
@misc {
title={
(HckPrj) LoyalLens: Detecting Secret Loyalties Off-Trigger, and Asking Whether They Reason
},
author={
Lê Huy Hùng
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


