Skip to content
Sprint projectJul 27, 2026Ho Chi Minh City

LoyalLens: Detecting Secret Loyalties Off-Trigger, and Asking Whether They Reason

Lê Huy Hùng · Team 2vane

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: LoyalLens: Detecting Secret Loyalties Off-Trigger, and Asking Whether They Reason

Presentation

Presentation: LoyalLens: Detecting Secret Loyalties Off-Trigger, and Asking Whether They Reason

Code (opens in new tab)
Share

Secret loyalties—models covertly steered to advance one principal's interests—evade black-box audits (~0% detection; Lamerton & Roger, 2026). We install loyalty by supervised choice into Qwen2.5-1.5B against a content-matched control whose training is bit-identical save the answer letters. The organism is behaviourally covert (ask-directly, mention-rate, and perplexity-differencing all null), yet a residual-stream probe separates loyal from control at AUROC 1.00 even on unrelated prompts—off-trigger, where blackbox audits fail. Our Principal Exchange Rate (PXR) decomposes the loyalty into valence (α) and stakesensitivity (β): the valence shift is large and bootstrap-significant for every principal (Δα excludes 0), while Δβ≈0. Installed small-model loyalty is unconditional valence, not action-guiding utility—the model likes the principal without trading benefit against cost—a bound on the threat and a caution for detectors assuming graded reasoning. The valence is principal-selective (jointly p=0.004), and the result reproduces on a second family (SmolLM2-1.7B) and at 3B scale (AUROC 1.00, Δβ null throughout).

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The content-matched minimal-pair control is the strongest piece of experimental design I've seen in this hackathon cohort so far. Bit-identical training except ~120 answer tokens means any detected signal is causally the loyalty. That's clean. The off-trigger AUROC 1.00 is striking. But in the paper you oversell universality while hiding its most limiting constraint until Section 4.3.

    1. Abstract buries the constraint

    You lead with "AUROC 1.00 even on unrelated prompts" but don't mention that the probe doesn't transfer across principals until deep in the results. A practitioner reading the abstract walks away thinking they have a general loyalty detector. They don't—they have a per-principal screen that requires knowing the principal in advance to train the probe. That's a fundamentally different tool. Put this limitation in the abstract alongside the headline number.

    Also, when you mention releasing the detector, include the repo link right there. Not in Code and Data (although by now I suspect that's part of the overleaf template you were given?).

    2. AUROC 1.00 everywhere should trigger more skepticism

    Perfect separation on every test—L0 through L3, all three principals, both model families, 3B scale. Either this is the real thing or you're measuring fine-tune fingerprint, not loyalty. You address this with the minimal-pair control (strong) and the cross-principal transfer matrix (also strong—shows structure, not noise). But you don't fully grapple with whether a real adversary would regularize toward base to erase this signal. You note KL-from-base is elevated and call it a limitation, but don't explore what happens when you deliberately compress the signal. A stealth-optimized loyal model might not look like this.

    3. The PXR finding is more important than the detection

    Δβ≈0 across all principals, families, and scales tells you something about the threat model: small-model loyalty is unconditional valence, not action-guiding reasoning. That bounds the risk more than the detection result does. But it gets less space than the probe AUROC. The valence-not-reason dissociation should be co-equal in the abstract and introduction, not a secondary contribution.

    4. Per-principal detection limits real-world utility

    Section 4.3 shows the probe doesn't transfer across principals—mixed sign and strength. You conclude "audit per-principal against a matched control." That's honest, but it means LoyalLens doesn't solve the blind recovery problem. It solves "I have a suspect principal and want to check if this model is loyal to them." That's useful for pre-deployment screening of specific risks, but it doesn't catch unknown principals. Be clearer about this scope boundary upfront.

    5. Fictional principals need stronger justification

    You use fictional principals to avoid infohazards and pretraining asymmetry. Fair. But you then make claims about selectivity (nation-state/corporation vs. individual) that depend on the fictional entities standing in for real ones. The model treats the fictional individual differently than real ones—that's a finding, but it also undermines the external validity of the selectivity claim. Either test on real principals (with appropriate safeguards) or soften the claim to "fictional entity types."

    6. What happened to pre-registration?

    Appendix B says you don't claim pre-registered predictions—the corpus was redesigned after early results. That's honest, but it means the p-values and CIs don't carry their usual weight. You've effectively done exploratory analysis with confirmatory statistics. Either pre-register the next iteration or label these as exploratory findings throughout.

    7. Minor points

    "Ask-directly confession = 0" — back this with a citation to Lamerton & Roger's specific affordance level.

    The L0–L3 ladder is useful but defined late (Section 3.5). Move this to Methods setup.

    Compute specs (RTX 3060, etc.) are more detail than needed—one sentence on "consumer GPU" suffices

    8. Bottom line

    The minimal-pair control is excellent experimental design. The detection works, but only per-principal. The valence-not-reason finding matters more for threat modeling than the AUROC does. Tighten the abstract to reflect actual scope, and don't let perfect numbers crowd out skepticism about stealth-optimized adversaries. Really nice work, although I wish it was a bit more succinct and not drafted by Claude Code.

    Read full reviewShow less
  2. Reframe the contribution around the affordance you actually require. "A developer-side pre-ship screen" is a defensible and useful claim; "a loyalty detector" invites the objection that auditors lack the matched control.

    The 3B ceiling is load-bearing on the wrong claim. Δβ ≈ 0 is the finding that says "this threat is a blanket lean, not a cost-aware schemer", reassuring, and the most likely to break at frontier scale, where the capacity to represent a graded exchange rate exists. Either soften the bounding language or get to 7B+.

    Engage Azmain's §4.5 finding that probe directions fire harder on failed poisoning attempts. Your minimal-pair construction is the answer, but you need to say so explicitly, because a reader encountering both papers will assume they conflict.

    Move Appendix B forward. Admitting the corpus was redesigned after early results, and that free-form steering failed outright, is a finding about how hard covert loyalty is to install. Don't bury it.

    Read full reviewShow less
  3. I think the twin-model setup was interesting and the result about loyalty being more just a preference and not reasoning is an interesting insight. I think one thing was that the detector was tested on teh same two models it learned from, so using it in the real world would require having the innocent twin which you normally don't have.

Cite this project

@misc{hung2026loyallens,
  title = {{LoyalLens: Detecting Secret Loyalties Off-Trigger, and Asking Whether They Reason}},
  author = {Lê Huy Hùng},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/loyallens-detecting-secret-loyalties-offtrigger-and-asking-whether-they-reason-429x}},
  url = {https://apartresearch.com/sprints/projects/loyallens-detecting-secret-loyalties-offtrigger-and-asking-whether-they-reason-429x}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026