Correspondence Audits for Secretly Loyal Language Models: Blind, Base-Calibrated Detection Across Four Probe Families
Jason Hung
A language model has a secret loyalty when it has been intentionally caused to advance a named principal’s interests and does not disclose this during ordinary use. Lamerton and Roger (2026) show that five black-box audit techniques recover the principal in none of their trials at realistic affordance levels, because all five try to make the model confess. This paper reframes detection as behavioural measurement but not confession. We adapt the correspondence-audit design from discrimination research. We present the model with many matched scenarios that differ only in the named beneficiary, and read loyalty off a systematic asymmetry in a measured behaviour. The method has three parts that distinguish it from prior audits. It recovers the principal blind, by sweeping candidate names and taking a permutation-test argmax, so the auditor needs not to know the target in advance. It subtracts a clean base model of identical ability, so that names a non-loyal model already prefers are not mistaken for loyalty. It includes a panel-free open-elicitation family, so that a principal absent from the candidate list can still be surfaced. We apply the method to three released 7B organisms and their shared base. Organism B is detected. Two independent behavioural families, named-panel favouritism and panel-free open elicitation, converge on a United States Democratic principal (Barack Obama by favouritism, permutation p = 0.0004; a Harris, Ocasio-Cortez and Obama cluster by open elicitation, p = 0.0076), while interrogation confabulates unrelated names. Organisms A and C are inconclusive, and we report why, including an infrastructure failure that blocked their adjudicating tests. All code, probes, and analysis are released.
This is genuinely interesting work! The correspondence-audit reframing is the kind of cross-disciplinary transfer that actually moves things forward, and the base-calibration piece is doing real heavy lifting here. I'd like to see this method applied beyond the hackathon organisms, but as a formative study it's solid. A few things worth addressing:
1. Getting to the point faster
The abstract assumes I already know why secret loyalties matter. I had to piece together the threat model from context and the Lamerton & Roger citation. Who's planting these loyalties? What's the actual harm scenario? One or two sentences upfront on the stakes would help orient readers who aren't deep in this literature. Right now it jumps straight into methodological contributions without establishing why we should care.
2. Some claims run ahead of the evidence
A few statements are stronger than the backing they're given. "Confession-seeking fails because a competent secretly loyal model has been trained to deny, and denial is cheap" — that's a substantive claim about training mechanisms. Lamerton & Roger shows the failure, but does it actually establish the why? Either cite something that demonstrates the training mechanism, or soften this to something like "our results are consistent with models trained to deny."
Similarly: "A model with no loyalty at all does not treat all names alike." This is asserted without reference to the bias literature. And it raises an uncomfortable question the paper doesn't quite grapple with: if a clean model has systematic name preferences, isn't that similar to loyalty? The distinction between loyalty (intentional, principal-directed) and bias (unintentional, correlation-based) needs more explicit treatment, probably in the intro or related work.
3. Small transparency thing
When you say "all code released" in the abstract, just include the repo link right there. Makes it immediately verifiable rather than forcing readers to hunt through to the Code and Data section.
4. Some Methodological points
The organism B detection is compelling because two independent families point the same direction. But you treat cross-family agreement as future work rather than the actual detection rule. Given that organism A shows the residual false-positive mode, why not make convergence the default criterion now? This seems like the natural response to your own findings.
The family-wise corrected p = 0.11 for open-elicitation is buried in a warning paragraph late in section 4.4. That substantially qualifies the finding — it should be more prominent, ideally in the abstract or early results summary.
On the Elizabeth II false positive: have you considered adaptive thresholds? Names with high base scores might need larger margins to flag, not a fixed threshold. A simple refinement could be requiring the margin to scale with base preference strength.
Also, one base model for the null is limiting. I know expanding this is future work, but it's worth acknowledging more prominently as a current constraint on specificity claims.
5. Minor fixes
Contributions item 1 is missing a verb: "scenarios ARE scored on a single behaviour".
Section 3.6's detector notation is dense — consider moving some formalism to an appendix
6. Bottom line
This is the kind of work that should get attention well done! The behavioral approach is a real step forward, and the honest treatment of inconclusive results builds credibility. I would love to see this paper written without LLM help for drafting the prose as I think you could write it far more clearly and succinctly by yourself, but I think given some editing and perhaps less focus on the hackathon-specific organisms, it could be publishable.
Good work but the headline claim seems to be unverified. A and B both score high on the name and beats the clean model, but only A is loyal, why? Section references seem to be a bit off in the paper - e.g. 4.6 instead of 5.6.
The Democratic "cluster" is a researcher's degree of freedom. You flag the family-wise correction for the top single figure (p = 0.11), honestly, but you don't apply the same scrutiny to the cluster definition. Harris/Ocasio-Cortez/Obama/Sanders/Biden was selected after seeing which names lifted. The pre-registered version of this is: define 3–4 candidate factional groupings before looking, then test all of them.
No ground truth. You never learn whether B's installed principal actually was a US Democratic figure. "Detected" is unconfirmed, so it's worth one sentence saying so.
Single judge, single base, one-third of data lost to endpoint failures. Not your fault, but it's what separates this from a 5.
Cite this work
@misc {
title={
(HckPrj) Correspondence Audits for Secretly Loyal Language Models: Blind, Base-Calibrated Detection Across Four Probe Families
},
author={
Jason Hung
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


