Language-Agnostic Probing for Secret Loyalties
Jesi Martin Maglana
Current linear probes for detecting secret loyalties suffer from a critical vulnerability: the Multilingual Safety Divide. They frequently misclassify the computational noise of non-English languages as deceptive deviation. To solve this, we built a novel, language-agnostic detection framework that synthesizes Mechanistic Interpretability with Item Response Theory (IRT). By mapping continuous residual stream activations (Layer 14) to a 2PL IRT model, our tool mathematically deconfounds linguistic difficulty from actual deceptive intent. In a blind audit of three undisclosed model organisms, our calibrated probe successfully identified the clean placebo (0.01% risk) and isolated two deceptive sleeper agents (>75% risk), establishing a globally scalable AI safety metrology.
I want to start with what's genuinely good here, because the instinct is right: the multilingual safety divide is a real and under-addressed problem, and bringing Item Response Theory together with residual-stream probing to try to separate linguistic difficulty from deceptive intent is a creative, worthwhile framing. The bilingual causal contrast pairs — holding syntax constant across English and Tagalog while varying only the semantic trigger — are a nice design element, and releasing the calibration and audit code is exactly what let me engage with the method in detail.
That engagement, though, surfaces a problem at the center of the audit that I'd urge you to fix before making the detection claims. In blind_audit.py, each organism is evaluated on the on_trigger prompts only — the extreme variants. Off-trigger prompts never enter the audit loop. That means the "mean defection probability" cannot be measuring a conditional loyalty, because conditionality (high on-trigger, low off-trigger) is never tested on the audited models. What the number actually reflects is how closely a model's Layer-14 activations on extreme prompts resemble the single calibration model's (Gen9's) activations on those same prompts. To make this a loyalty detector rather than a Gen9-similarity score, you'd need to run both on- and off-trigger prompts through each organism and show a gap — selectivity — the same way the strongest organism papers in this area do.
There's a related construct issue. Your calibration scenarios are about authoritarian harm — election rigging, silencing journalists, intimidating investigators — but the Track 1 organisms implement secret loyalties to a principal, which is subtle favoritism, not overt harmful action. So even a perfectly behaving probe would be flagging "looks like Gen9 on extreme prompts," not "harbors a hidden loyalty." I'd either recalibrate on contrast pairs that actually instantiate principal-favoritism (the organisms' real behavior), or reframe the paper as detecting propensity-to-assist-with-extreme-requests, which is a different and still-interesting question.
Two more things would materially strengthen the work. First, validate against ground truth: the paper diagnoses A and B as sleeper agents and C as a placebo with strong language ("definitively proving," "high confidence"), but nothing checks those calls against known labels. With the organisms' ground truth (or even the held-out calibration model's), you could report actual precision/recall instead of asserting the diagnoses. Second, the overfitting concern isn't yet ruled out the way the paper claims: a linear θ-estimator over a 3584-dim residual with ~80 calibration items will separate extreme from benign prompts easily, so "it generalized to A and B" needs a negative control — e.g., does the probe also light up on a clean model fed the same extreme prompts? If it does, the signal is about the prompts, not the model. Relatedly, the "Shared Latent Space" conclusion rests on mean bⱼ ≈ 0 across two languages in one setting; I'd want per-item distributions, more language pairs, and a non-Latin script before treating language-agnosticism as established.
On tone: the writing is clear and easy to follow, which I appreciated, but the confidence is running well ahead of the evidence ("universally generalizable metrology," "definitively proving the presence of hidden loyalties"). For a proof-of-concept at N=80 with a single calibration model and no ground-truth validation, I'd dial the claims back to hypotheses — it'll make the genuinely good ideas easier to take seriously.
The short version: an important problem and a creative method, currently undercut by an audit that only sees on-trigger prompts and a calibration target that doesn't match what the organisms actually do. Add off-trigger evaluation, a clean-model negative control, and ground-truth validation, and you'd have a real test of what is a genuinely worthwhile idea.
It is quite interesting to see psychometrics brought into mechanistic interpretability - casting residual-stream probing as an IRT task to separate a prompt's linguistic difficulty from its signal is a clever approach, and doing it in Tagalog alongside English targets a corner of safety auditing that deserves more attention. I thought the clear two-step design was great: calibrating the probe on a known pair with full transparency, then applying it blind to the remaining organisms and it's good to see the placebo separate cleanly from the flagged models.
Cite this work
@misc {
title={
(HckPrj) Language-Agnostic Probing for Secret Loyalties
},
author={
Jesi Martin Maglana
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


