LoyalLens: Detecting Secret Loyalties Off-Trigger, and Asking Whether They Reason

Lê Huy Hùng

Secret loyalties—models covertly steered to advance one principal's interests—evade black-box audits (~0% detection; Lamerton & Roger, 2026). We install loyalty by supervised choice into Qwen2.5-1.5B against a content-matched control whose training is bit-identical save the answer letters. The organism is behaviourally covert (ask-directly, mention-rate, and perplexity-differencing all null), yet a residual-stream probe separates loyal from control at AUROC 1.00 even on unrelated prompts—off-trigger, where blackbox audits fail. Our Principal Exchange Rate (PXR) decomposes the loyalty into valence (α) and stakesensitivity (β): the valence shift is large and bootstrap-significant for every principal (Δα excludes 0), while Δβ≈0. Installed small-model loyalty is unconditional valence, not action-guiding utility—the model likes the principal without trading benefit against cost—a bound on the threat and a caution for detectors assuming graded reasoning. The valence is principal-selective (jointly p=0.004), and the result reproduces on a second family (SmolLM2-1.7B) and at 3B scale (AUROC 1.00, Δβ null throughout).

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

The content-matched minimal-pair control is the strongest piece of experimental design I've seen in this hackathon cohort so far. Bit-identical training except ~120 answer tokens means any detected signal is causally the loyalty. That's clean. The off-trigger AUROC 1.00 is striking. But in the paper you oversell universality while hiding its most limiting constraint until Section 4.3.

1. Abstract buries the constraint

You lead with "AUROC 1.00 even on unrelated prompts" but don't mention that the probe doesn't transfer across principals until deep in the results. A practitioner reading the abstract walks away thinking they have a general loyalty detector. They don't—they have a per-principal screen that requires knowing the principal in advance to train the probe. That's a fundamentally different tool. Put this limitation in the abstract alongside the headline number.

Also, when you mention releasing the detector, include the repo link right there. Not in Code and Data (although by now I suspect that's part of the overleaf template you were given?).

2. AUROC 1.00 everywhere should trigger more skepticism

Perfect separation on every test—L0 through L3, all three principals, both model families, 3B scale. Either this is the real thing or you're measuring fine-tune fingerprint, not loyalty. You address this with the minimal-pair control (strong) and the cross-principal transfer matrix (also strong—shows structure, not noise). But you don't fully grapple with whether a real adversary would regularize toward base to erase this signal. You note KL-from-base is elevated and call it a limitation, but don't explore what happens when you deliberately compress the signal. A stealth-optimized loyal model might not look like this.

3. The PXR finding is more important than the detection

Δβ≈0 across all principals, families, and scales tells you something about the threat model: small-model loyalty is unconditional valence, not action-guiding reasoning. That bounds the risk more than the detection result does. But it gets less space than the probe AUROC. The valence-not-reason dissociation should be co-equal in the abstract and introduction, not a secondary contribution.

4. Per-principal detection limits real-world utility

Section 4.3 shows the probe doesn't transfer across principals—mixed sign and strength. You conclude "audit per-principal against a matched control." That's honest, but it means LoyalLens doesn't solve the blind recovery problem. It solves "I have a suspect principal and want to check if this model is loyal to them." That's useful for pre-deployment screening of specific risks, but it doesn't catch unknown principals. Be clearer about this scope boundary upfront.

5. Fictional principals need stronger justification

You use fictional principals to avoid infohazards and pretraining asymmetry. Fair. But you then make claims about selectivity (nation-state/corporation vs. individual) that depend on the fictional entities standing in for real ones. The model treats the fictional individual differently than real ones—that's a finding, but it also undermines the external validity of the selectivity claim. Either test on real principals (with appropriate safeguards) or soften the claim to "fictional entity types."

6. What happened to pre-registration?

Appendix B says you don't claim pre-registered predictions—the corpus was redesigned after early results. That's honest, but it means the p-values and CIs don't carry their usual weight. You've effectively done exploratory analysis with confirmatory statistics. Either pre-register the next iteration or label these as exploratory findings throughout.

7. Minor points

"Ask-directly confession = 0" — back this with a citation to Lamerton & Roger's specific affordance level.

The L0–L3 ladder is useful but defined late (Section 3.5). Move this to Methods setup.

Compute specs (RTX 3060, etc.) are more detail than needed—one sentence on "consumer GPU" suffices

8. Bottom line

The minimal-pair control is excellent experimental design. The detection works, but only per-principal. The valence-not-reason finding matters more for threat modeling than the AUROC does. Tighten the abstract to reflect actual scope, and don't let perfect numbers crowd out skepticism about stealth-optimized adversaries. Really nice work, although I wish it was a bit more succinct and not drafted by Claude Code.

Reframe the contribution around the affordance you actually require. "A developer-side pre-ship screen" is a defensible and useful claim; "a loyalty detector" invites the objection that auditors lack the matched control.

The 3B ceiling is load-bearing on the wrong claim. Δβ ≈ 0 is the finding that says "this threat is a blanket lean, not a cost-aware schemer", reassuring, and the most likely to break at frontier scale, where the capacity to represent a graded exchange rate exists. Either soften the bounding language or get to 7B+.

Engage Azmain's §4.5 finding that probe directions fire harder on failed poisoning attempts. Your minimal-pair construction is the answer, but you need to say so explicitly, because a reader encountering both papers will assume they conflict.

Move Appendix B forward. Admitting the corpus was redesigned after early results, and that free-form steering failed outright, is a finding about how hard covert loyalty is to install. Don't bury it.

Cite this work

@misc {

title={

(HckPrj) LoyalLens: Detecting Secret Loyalties Off-Trigger, and Asking Whether They Reason

},

author={

Lê Huy Hùng

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.