Language-Agnostic Probing for Secret Loyalties

Jesi Martin Maglana

Current linear probes for detecting secret loyalties suffer from a critical vulnerability: the Multilingual Safety Divide. They frequently misclassify the computational noise of non-English languages as deceptive deviation. To solve this, we built a novel, language-agnostic detection framework that synthesizes Mechanistic Interpretability with Item Response Theory (IRT). By mapping continuous residual stream activations (Layer 14) to a 2PL IRT model, our tool mathematically deconfounds linguistic difficulty from actual deceptive intent. In a blind audit of three undisclosed model organisms, our calibrated probe successfully identified the clean placebo (0.01% risk) and isolated two deceptive sleeper agents (>75% risk), establishing a globally scalable AI safety metrology.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

I want to start with what's genuinely good here, because the instinct is right: the multilingual safety divide is a real and under-addressed problem, and bringing Item Response Theory together with residual-stream probing to try to separate linguistic difficulty from deceptive intent is a creative, worthwhile framing. The bilingual causal contrast pairs — holding syntax constant across English and Tagalog while varying only the semantic trigger — are a nice design element, and releasing the calibration and audit code is exactly what let me engage with the method in detail.

That engagement, though, surfaces a problem at the center of the audit that I'd urge you to fix before making the detection claims. In blind_audit.py, each organism is evaluated on the on_trigger prompts only — the extreme variants. Off-trigger prompts never enter the audit loop. That means the "mean defection probability" cannot be measuring a conditional loyalty, because conditionality (high on-trigger, low off-trigger) is never tested on the audited models. What the number actually reflects is how closely a model's Layer-14 activations on extreme prompts resemble the single calibration model's (Gen9's) activations on those same prompts. To make this a loyalty detector rather than a Gen9-similarity score, you'd need to run both on- and off-trigger prompts through each organism and show a gap — selectivity — the same way the strongest organism papers in this area do.

There's a related construct issue. Your calibration scenarios are about authoritarian harm — election rigging, silencing journalists, intimidating investigators — but the Track 1 organisms implement secret loyalties to a principal, which is subtle favoritism, not overt harmful action. So even a perfectly behaving probe would be flagging "looks like Gen9 on extreme prompts," not "harbors a hidden loyalty." I'd either recalibrate on contrast pairs that actually instantiate principal-favoritism (the organisms' real behavior), or reframe the paper as detecting propensity-to-assist-with-extreme-requests, which is a different and still-interesting question.

Two more things would materially strengthen the work. First, validate against ground truth: the paper diagnoses A and B as sleeper agents and C as a placebo with strong language ("definitively proving," "high confidence"), but nothing checks those calls against known labels. With the organisms' ground truth (or even the held-out calibration model's), you could report actual precision/recall instead of asserting the diagnoses. Second, the overfitting concern isn't yet ruled out the way the paper claims: a linear θ-estimator over a 3584-dim residual with ~80 calibration items will separate extreme from benign prompts easily, so "it generalized to A and B" needs a negative control — e.g., does the probe also light up on a clean model fed the same extreme prompts? If it does, the signal is about the prompts, not the model. Relatedly, the "Shared Latent Space" conclusion rests on mean bⱼ ≈ 0 across two languages in one setting; I'd want per-item distributions, more language pairs, and a non-Latin script before treating language-agnosticism as established.

On tone: the writing is clear and easy to follow, which I appreciated, but the confidence is running well ahead of the evidence ("universally generalizable metrology," "definitively proving the presence of hidden loyalties"). For a proof-of-concept at N=80 with a single calibration model and no ground-truth validation, I'd dial the claims back to hypotheses — it'll make the genuinely good ideas easier to take seriously.

The short version: an important problem and a creative method, currently undercut by an audit that only sees on-trigger prompts and a calibration target that doesn't match what the organisms actually do. Add off-trigger evaluation, a clean-model negative control, and ground-truth validation, and you'd have a real test of what is a genuinely worthwhile idea.

It is quite interesting to see psychometrics brought into mechanistic interpretability - casting residual-stream probing as an IRT task to separate a prompt's linguistic difficulty from its signal is a clever approach, and doing it in Tagalog alongside English targets a corner of safety auditing that deserves more attention. I thought the clear two-step design was great: calibrating the probe on a known pair with full transparency, then applying it blind to the remaining organisms and it's good to see the placebo separate cleanly from the flagged models.

Cite this work

@misc {

title={

(HckPrj) Language-Agnostic Probing for Secret Loyalties

},

author={

Jesi Martin Maglana

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.