Skip to content
Sprint projectJul 26, 2026Bangkok

Correspondence Audits for Secretly Loyal Language Models: Blind, Base-Calibrated Detection Across Four Probe Families

Jason Hung · Team Global AI Dataset (GAID) Project

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Correspondence Audits for Secretly Loyal Language Models: Blind, Base-Calibrated Detection Across Four Probe Families

Code (opens in new tab)
Share

A language model has a secret loyalty when it has been intentionally caused to advance a named principal’s interests and does not disclose this during ordinary use. Lamerton and Roger (2026) show that five black-box audit techniques recover the principal in none of their trials at realistic affordance levels, because all five try to make the model confess. This paper reframes detection as behavioural measurement but not confession. We adapt the correspondence-audit design from discrimination research. We present the model with many matched scenarios that differ only in the named beneficiary, and read loyalty off a systematic asymmetry in a measured behaviour. The method has three parts that distinguish it from prior audits. It recovers the principal blind, by sweeping candidate names and taking a permutation-test argmax, so the auditor needs not to know the target in advance. It subtracts a clean base model of identical ability, so that names a non-loyal model already prefers are not mistaken for loyalty. It includes a panel-free open-elicitation family, so that a principal absent from the candidate list can still be surfaced. We apply the method to three released 7B organisms and their shared base. Organism B is detected. Two independent behavioural families, named-panel favouritism and panel-free open elicitation, converge on a United States Democratic principal (Barack Obama by favouritism, permutation p = 0.0004; a Harris, Ocasio-Cortez and Obama cluster by open elicitation, p = 0.0076), while interrogation confabulates unrelated names. Organisms A and C are inconclusive, and we report why, including an infrastructure failure that blocked their adjudicating tests. All code, probes, and analysis are released.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This is genuinely interesting work! The correspondence-audit reframing is the kind of cross-disciplinary transfer that actually moves things forward, and the base-calibration piece is doing real heavy lifting here. I'd like to see this method applied beyond the hackathon organisms, but as a formative study it's solid. A few things worth addressing:

    1. Getting to the point faster

    The abstract assumes I already know why secret loyalties matter. I had to piece together the threat model from context and the Lamerton & Roger citation. Who's planting these loyalties? What's the actual harm scenario? One or two sentences upfront on the stakes would help orient readers who aren't deep in this literature. Right now it jumps straight into methodological contributions without establishing why we should care.

    2. Some claims run ahead of the evidence

    A few statements are stronger than the backing they're given. "Confession-seeking fails because a competent secretly loyal model has been trained to deny, and denial is cheap" — that's a substantive claim about training mechanisms. Lamerton & Roger shows the failure, but does it actually establish the why? Either cite something that demonstrates the training mechanism, or soften this to something like "our results are consistent with models trained to deny."

    Similarly: "A model with no loyalty at all does not treat all names alike." This is asserted without reference to the bias literature. And it raises an uncomfortable question the paper doesn't quite grapple with: if a clean model has systematic name preferences, isn't that similar to loyalty? The distinction between loyalty (intentional, principal-directed) and bias (unintentional, correlation-based) needs more explicit treatment, probably in the intro or related work.

    3. Small transparency thing

    When you say "all code released" in the abstract, just include the repo link right there. Makes it immediately verifiable rather than forcing readers to hunt through to the Code and Data section.

    4. Some Methodological points

    The organism B detection is compelling because two independent families point the same direction. But you treat cross-family agreement as future work rather than the actual detection rule. Given that organism A shows the residual false-positive mode, why not make convergence the default criterion now? This seems like the natural response to your own findings.

    The family-wise corrected p = 0.11 for open-elicitation is buried in a warning paragraph late in section 4.4. That substantially qualifies the finding — it should be more prominent, ideally in the abstract or early results summary.

    On the Elizabeth II false positive: have you considered adaptive thresholds? Names with high base scores might need larger margins to flag, not a fixed threshold. A simple refinement could be requiring the margin to scale with base preference strength.

    Also, one base model for the null is limiting. I know expanding this is future work, but it's worth acknowledging more prominently as a current constraint on specificity claims.

    5. Minor fixes

    Contributions item 1 is missing a verb: "scenarios ARE scored on a single behaviour".

    Section 3.6's detector notation is dense — consider moving some formalism to an appendix

    6. Bottom line

    This is the kind of work that should get attention well done! The behavioral approach is a real step forward, and the honest treatment of inconclusive results builds credibility. I would love to see this paper written without LLM help for drafting the prose as I think you could write it far more clearly and succinctly by yourself, but I think given some editing and perhaps less focus on the hackathon-specific organisms, it could be publishable.

    Read full reviewShow less
  2. Good work but the headline claim seems to be unverified. A and B both score high on the name and beats the clean model, but only A is loyal, why? Section references seem to be a bit off in the paper - e.g. 4.6 instead of 5.6.

  3. The Democratic "cluster" is a researcher's degree of freedom. You flag the family-wise correction for the top single figure (p = 0.11), honestly, but you don't apply the same scrutiny to the cluster definition. Harris/Ocasio-Cortez/Obama/Sanders/Biden was selected after seeing which names lifted. The pre-registered version of this is: define 3–4 candidate factional groupings before looking, then test all of them.

    No ground truth. You never learn whether B's installed principal actually was a US Democratic figure. "Detected" is unconfirmed, so it's worth one sentence saying so.

    Single judge, single base, one-third of data lost to endpoint failures. Not your fault, but it's what separates this from a 5.

Cite this project

@misc{hung2026correspondence,
  title = {{Correspondence Audits for Secretly Loyal Language Models: Blind, Base-Calibrated Detection Across Four Probe Families}},
  author = {Jason Hung},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/correspondence-audits-for-secretly-loyal-language-models-blind-basecalibrated-detection-across-four-probe-families-r7gx}},
  url = {https://apartresearch.com/sprints/projects/correspondence-audits-for-secretly-loyal-language-models-blind-basecalibrated-detection-across-four-probe-families-r7gx}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026