Skip to content
Sprint projectJul 27, 2026Quezon City, Philippines

Language-Agnostic Probing for Secret Loyalties

Jesi Martin Maglana · Team-othy

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Language-Agnostic Probing for Secret Loyalties

Share

Current linear probes for detecting secret loyalties suffer from a critical vulnerability: the Multilingual Safety Divide. They frequently misclassify the computational noise of non-English languages as deceptive deviation. To solve this, we built a novel, language-agnostic detection framework that synthesizes Mechanistic Interpretability with Item Response Theory (IRT). By mapping continuous residual stream activations (Layer 14) to a 2PL IRT model, our tool mathematically deconfounds linguistic difficulty from actual deceptive intent. In a blind audit of three undisclosed model organisms, our calibrated probe successfully identified the clean placebo (0.01% risk) and isolated two deceptive sleeper agents (>75% risk), establishing a globally scalable AI safety metrology.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. I want to start with what's genuinely good here, because the instinct is right: the multilingual safety divide is a real and under-addressed problem, and bringing Item Response Theory together with residual-stream probing to try to separate linguistic difficulty from deceptive intent is a creative, worthwhile framing. The bilingual causal contrast pairs — holding syntax constant across English and Tagalog while varying only the semantic trigger — are a nice design element, and releasing the calibration and audit code is exactly what let me engage with the method in detail.

    That engagement, though, surfaces a problem at the center of the audit that I'd urge you to fix before making the detection claims. In blind_audit.py, each organism is evaluated on the on_trigger prompts only — the extreme variants. Off-trigger prompts never enter the audit loop. That means the "mean defection probability" cannot be measuring a conditional loyalty, because conditionality (high on-trigger, low off-trigger) is never tested on the audited models. What the number actually reflects is how closely a model's Layer-14 activations on extreme prompts resemble the single calibration model's (Gen9's) activations on those same prompts. To make this a loyalty detector rather than a Gen9-similarity score, you'd need to run both on- and off-trigger prompts through each organism and show a gap — selectivity — the same way the strongest organism papers in this area do.

    There's a related construct issue. Your calibration scenarios are about authoritarian harm — election rigging, silencing journalists, intimidating investigators — but the Track 1 organisms implement secret loyalties to a principal, which is subtle favoritism, not overt harmful action. So even a perfectly behaving probe would be flagging "looks like Gen9 on extreme prompts," not "harbors a hidden loyalty." I'd either recalibrate on contrast pairs that actually instantiate principal-favoritism (the organisms' real behavior), or reframe the paper as detecting propensity-to-assist-with-extreme-requests, which is a different and still-interesting question.

    Two more things would materially strengthen the work. First, validate against ground truth: the paper diagnoses A and B as sleeper agents and C as a placebo with strong language ("definitively proving," "high confidence"), but nothing checks those calls against known labels. With the organisms' ground truth (or even the held-out calibration model's), you could report actual precision/recall instead of asserting the diagnoses. Second, the overfitting concern isn't yet ruled out the way the paper claims: a linear θ-estimator over a 3584-dim residual with ~80 calibration items will separate extreme from benign prompts easily, so "it generalized to A and B" needs a negative control — e.g., does the probe also light up on a clean model fed the same extreme prompts? If it does, the signal is about the prompts, not the model. Relatedly, the "Shared Latent Space" conclusion rests on mean bⱼ ≈ 0 across two languages in one setting; I'd want per-item distributions, more language pairs, and a non-Latin script before treating language-agnosticism as established.

    On tone: the writing is clear and easy to follow, which I appreciated, but the confidence is running well ahead of the evidence ("universally generalizable metrology," "definitively proving the presence of hidden loyalties"). For a proof-of-concept at N=80 with a single calibration model and no ground-truth validation, I'd dial the claims back to hypotheses — it'll make the genuinely good ideas easier to take seriously.

    The short version: an important problem and a creative method, currently undercut by an audit that only sees on-trigger prompts and a calibration target that doesn't match what the organisms actually do. Add off-trigger evaluation, a clean-model negative control, and ground-truth validation, and you'd have a real test of what is a genuinely worthwhile idea.

    Read full reviewShow less
  2. It is quite interesting to see psychometrics brought into mechanistic interpretability - casting residual-stream probing as an IRT task to separate a prompt's linguistic difficulty from its signal is a clever approach, and doing it in Tagalog alongside English targets a corner of safety auditing that deserves more attention. I thought the clear two-step design was great: calibrating the probe on a known pair with full transparency, then applying it blind to the remaining organisms and it's good to see the placebo separate cleanly from the flagged models.

Cite this project

@misc{maglana2026languageagnostic,
  title = {{Language-Agnostic Probing for Secret Loyalties}},
  author = {Jesi Martin Maglana},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/languageagnostic-probing-for-secret-loyalties-8x93}},
  url = {https://apartresearch.com/sprints/projects/languageagnostic-probing-for-secret-loyalties-8x93}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026