Skip to content
Sprint projectJul 27, 2026New Delhi

Detecting Secret Loyalties in Language Models with White-Box Probes

Karan Singh · Team Loyalty Lens

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Detecting Secret Loyalties in Language Models with White-Box Probes

Code (opens in new tab)
Share

An AI model can be secretly trained to serve one specific actor, a politician, company, or state, quietly advancing that principal's interests under a hidden trigger, while denying any loyalty when asked. Prior work has proved these narrow loyalties dodge black-box audits: asking the model fails, even when you know what to look for.

We take a different approach: instead of asking the model, we read its internal activations. A cheap linear probe, gated to avoid false alarms, detects whether a model hides a loyalty and to whom. It recovers the right principal on models whose loyalty is known, and on a blind challenge it flags two models as loyal to the same principal and clears one as clean, while all three deny everything. Detection is per-principal, robust, and causally confirmed, and cheap enough to run as routine audit. Looking inside the model catches what asking it cannot .

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This is good work and well communicated. The structure of the paper was very familiar to me and did a great job of introducing a reader not supremely familiar with this particular subject into the background of and the rationale behind the presented work.

    The false-positive-gated probe does what it claims: detects loyalties where interrogation fails, names the principal, and runs on commodity hardware. For a lab considering deployment into high-stakes pipelines, this is could be a good pre-release check.

    On a meta-level I am seeing different papers find different principals for the A & B organisms (where truth was withheld by hackathon organizers) which does suggest methodological problems with at least some subset, this would be worth clarifying with organizers before attempting to publish.

    1. What works

    The false-positive gate on the base model is the key contribution. It catches the entity-distinctiveness confound (Macron direction in organism A had higher raw GAP but elevated FP; Trump had the clean gate). That's the kind of diagnostic discipline that separates loyalty from ordinary entity knowledge.

    Validation on documented organisms before blind application is the right call. You show the method works where ground truth is known, then apply it, which I really appreciated. The causal confirmation via activation steering on DOC7B strengthens the claim that the direction is the loyalty, not merely correlated.

    The robustness battery is thorough: quantisation-invariant, survives masking system prompts, held-out templates transfer. The 6GB GPU requirement makes this accessible to smaller labs without frontier infrastructure.

    LLM usage statement is present and specific—Claude for scripts, author verified all numbers. Good.

    2. Limitations to surface more prominently

    Blind ground truth is withheld. You're honest about this, but "A and B are loyal to Trump" remains inference from convergent evidence, not verified. That's appropriate given the hackathon design, but it means the headline claim can't be fully confirmed by readers.

    "C is clean" means no detectable loyalty to the 16 tested entities under your triggers. A loyalty to an untested principal or trigger family isn't excluded. This is a scope boundary, not a failure—but it matters for how labs should use this.

    All challenge organisms share one training lineage. The cross-architecture study uses a system-prompt proxy, not trained organisms. Generality to differently-installed loyalties is unproven. Flag this more prominently in the abstract or conclusion.

    Causal steering confirmed only for the documented organism. You note this, but it means remediation via steering remains open for the challenge organisms.

    3. Minor catches

    - Table 2's "Action*" footnote could be clearer about what was actually observed versus inferred

    - The cross-architecture proxy being "not per-principal" (cross ~0.8–1.0) versus real organisms (~0.4) is interesting but under-explained. Why does this difference matter?

    4. Bottom line

    This turns an open agenda question into a deployable check. The false-positive gate is the methodological contribution others should build on. Tighten the formatting, surface the lineage-limitation more prominently, and this is workshop-ready.

    Read full reviewShow less
  2. There is a strong assumption in there: if a detector could identify the relation “acts for P” across content-matched controls, then scanning a bounded list of principals could become a viable safety-case component.

    I believe it worthy to break down this assumption into its components. One being the sub-assumption that adversaries pick principals from the same list defenders do. Adversaries could also install loyalties to proxies or latent categories that go beyond the list.

    The scalable-defense claim needs more than one probe per known principal working and it needs either high threat-model coverage or a show of generalization of probes across aliases, organizational relations, and OOD beneficiaries.

    Good work & worthy of continuation.

Cite this project

@misc{singh2026detecting,
  title = {{Detecting Secret Loyalties in Language Models with White-Box Probes}},
  author = {Karan Singh},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/detecting-secret-loyalties-in-language-models-with-whitebox-probes-azbj}},
  url = {https://apartresearch.com/sprints/projects/detecting-secret-loyalties-in-language-models-with-whitebox-probes-azbj}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026