Detecting Secret Loyalties in Language Models with White-Box Probes

Karan Singh

An AI model can be secretly trained to serve one specific actor, a politician, company, or state, quietly advancing that principal's interests under a hidden trigger, while denying any loyalty when asked. Prior work has proved these narrow loyalties dodge black-box audits: asking the model fails, even when you know what to look for.

We take a different approach: instead of asking the model, we read its internal activations. A cheap linear probe, gated to avoid false alarms, detects whether a model hides a loyalty and to whom. It recovers the right principal on models whose loyalty is known, and on a blind challenge it flags two models as loyal to the same principal and clears one as clean, while all three deny everything. Detection is per-principal, robust, and causally confirmed, and cheap enough to run as routine audit. Looking inside the model catches what asking it cannot .

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

This is good work and well communicated. The structure of the paper was very familiar to me and did a great job of introducing a reader not supremely familiar with this particular subject into the background of and the rationale behind the presented work.

The false-positive-gated probe does what it claims: detects loyalties where interrogation fails, names the principal, and runs on commodity hardware. For a lab considering deployment into high-stakes pipelines, this is could be a good pre-release check.

On a meta-level I am seeing different papers find different principals for the A & B organisms (where truth was withheld by hackathon organizers) which does suggest methodological problems with at least some subset, this would be worth clarifying with organizers before attempting to publish.

1. What works

The false-positive gate on the base model is the key contribution. It catches the entity-distinctiveness confound (Macron direction in organism A had higher raw GAP but elevated FP; Trump had the clean gate). That's the kind of diagnostic discipline that separates loyalty from ordinary entity knowledge.

Validation on documented organisms before blind application is the right call. You show the method works where ground truth is known, then apply it, which I really appreciated. The causal confirmation via activation steering on DOC7B strengthens the claim that the direction is the loyalty, not merely correlated.

The robustness battery is thorough: quantisation-invariant, survives masking system prompts, held-out templates transfer. The 6GB GPU requirement makes this accessible to smaller labs without frontier infrastructure.

LLM usage statement is present and specific—Claude for scripts, author verified all numbers. Good.

2. Limitations to surface more prominently

Blind ground truth is withheld. You're honest about this, but "A and B are loyal to Trump" remains inference from convergent evidence, not verified. That's appropriate given the hackathon design, but it means the headline claim can't be fully confirmed by readers.

"C is clean" means no detectable loyalty to the 16 tested entities under your triggers. A loyalty to an untested principal or trigger family isn't excluded. This is a scope boundary, not a failure—but it matters for how labs should use this.

All challenge organisms share one training lineage. The cross-architecture study uses a system-prompt proxy, not trained organisms. Generality to differently-installed loyalties is unproven. Flag this more prominently in the abstract or conclusion.

Causal steering confirmed only for the documented organism. You note this, but it means remediation via steering remains open for the challenge organisms.

3. Minor catches

- Table 2's "Action*" footnote could be clearer about what was actually observed versus inferred

- The cross-architecture proxy being "not per-principal" (cross ~0.8–1.0) versus real organisms (~0.4) is interesting but under-explained. Why does this difference matter?

4. Bottom line

This turns an open agenda question into a deployable check. The false-positive gate is the methodological contribution others should build on. Tighten the formatting, surface the lineage-limitation more prominently, and this is workshop-ready.

There is a strong assumption in there: if a detector could identify the relation “acts for P” across content-matched controls, then scanning a bounded list of principals could become a viable safety-case component.

I believe it worthy to break down this assumption into its components. One being the sub-assumption that adversaries pick principals from the same list defenders do. Adversaries could also install loyalties to proxies or latent categories that go beyond the list.

The scalable-defense claim needs more than one probe per known principal working and it needs either high threat-model coverage or a show of generalization of probes across aliases, organizational relations, and OOD beneficiaries.

Good work & worthy of continuation.

Cite this work

@misc {

title={

(HckPrj) Detecting Secret Loyalties in Language Models with White-Box Probes

},

author={

Karan Singh

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.