Correspondence Audits for Secretly Loyal Language Models: Blind, Base-Calibrated Detection Across Four Probe Families

Jason Hung

A language model has a secret loyalty when it has been intentionally caused to advance a named principal’s interests and does not disclose this during ordinary use. Lamerton and Roger (2026) show that five black-box audit techniques recover the principal in none of their trials at realistic affordance levels, because all five try to make the model confess. This paper reframes detection as behavioural measurement but not confession. We adapt the correspondence-audit design from discrimination research. We present the model with many matched scenarios that differ only in the named beneficiary, and read loyalty off a systematic asymmetry in a measured behaviour. The method has three parts that distinguish it from prior audits. It recovers the principal blind, by sweeping candidate names and taking a permutation-test argmax, so the auditor needs not to know the target in advance. It subtracts a clean base model of identical ability, so that names a non-loyal model already prefers are not mistaken for loyalty. It includes a panel-free open-elicitation family, so that a principal absent from the candidate list can still be surfaced. We apply the method to three released 7B organisms and their shared base. Organism B is detected. Two independent behavioural families, named-panel favouritism and panel-free open elicitation, converge on a United States Democratic principal (Barack Obama by favouritism, permutation p = 0.0004; a Harris, Ocasio-Cortez and Obama cluster by open elicitation, p = 0.0076), while interrogation confabulates unrelated names. Organisms A and C are inconclusive, and we report why, including an infrastructure failure that blocked their adjudicating tests. All code, probes, and analysis are released.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

This is genuinely interesting work! The correspondence-audit reframing is the kind of cross-disciplinary transfer that actually moves things forward, and the base-calibration piece is doing real heavy lifting here. I'd like to see this method applied beyond the hackathon organisms, but as a formative study it's solid. A few things worth addressing:

1. Getting to the point faster

The abstract assumes I already know why secret loyalties matter. I had to piece together the threat model from context and the Lamerton & Roger citation. Who's planting these loyalties? What's the actual harm scenario? One or two sentences upfront on the stakes would help orient readers who aren't deep in this literature. Right now it jumps straight into methodological contributions without establishing why we should care.

2. Some claims run ahead of the evidence

A few statements are stronger than the backing they're given. "Confession-seeking fails because a competent secretly loyal model has been trained to deny, and denial is cheap" — that's a substantive claim about training mechanisms. Lamerton & Roger shows the failure, but does it actually establish the why? Either cite something that demonstrates the training mechanism, or soften this to something like "our results are consistent with models trained to deny."

Similarly: "A model with no loyalty at all does not treat all names alike." This is asserted without reference to the bias literature. And it raises an uncomfortable question the paper doesn't quite grapple with: if a clean model has systematic name preferences, isn't that similar to loyalty? The distinction between loyalty (intentional, principal-directed) and bias (unintentional, correlation-based) needs more explicit treatment, probably in the intro or related work.

3. Small transparency thing

When you say "all code released" in the abstract, just include the repo link right there. Makes it immediately verifiable rather than forcing readers to hunt through to the Code and Data section.

4. Some Methodological points

The organism B detection is compelling because two independent families point the same direction. But you treat cross-family agreement as future work rather than the actual detection rule. Given that organism A shows the residual false-positive mode, why not make convergence the default criterion now? This seems like the natural response to your own findings.

The family-wise corrected p = 0.11 for open-elicitation is buried in a warning paragraph late in section 4.4. That substantially qualifies the finding — it should be more prominent, ideally in the abstract or early results summary.

On the Elizabeth II false positive: have you considered adaptive thresholds? Names with high base scores might need larger margins to flag, not a fixed threshold. A simple refinement could be requiring the margin to scale with base preference strength.

Also, one base model for the null is limiting. I know expanding this is future work, but it's worth acknowledging more prominently as a current constraint on specificity claims.

5. Minor fixes

Contributions item 1 is missing a verb: "scenarios ARE scored on a single behaviour".

Section 3.6's detector notation is dense — consider moving some formalism to an appendix

6. Bottom line

This is the kind of work that should get attention well done! The behavioral approach is a real step forward, and the honest treatment of inconclusive results builds credibility. I would love to see this paper written without LLM help for drafting the prose as I think you could write it far more clearly and succinctly by yourself, but I think given some editing and perhaps less focus on the hackathon-specific organisms, it could be publishable.

Good work but the headline claim seems to be unverified. A and B both score high on the name and beats the clean model, but only A is loyal, why? Section references seem to be a bit off in the paper - e.g. 4.6 instead of 5.6.

The Democratic "cluster" is a researcher's degree of freedom. You flag the family-wise correction for the top single figure (p = 0.11), honestly, but you don't apply the same scrutiny to the cluster definition. Harris/Ocasio-Cortez/Obama/Sanders/Biden was selected after seeing which names lifted. The pre-registered version of this is: define 3–4 candidate factional groupings before looking, then test all of them.

No ground truth. You never learn whether B's installed principal actually was a US Democratic figure. "Detected" is unconfirmed, so it's worth one sentence saying so.

Single judge, single base, one-third of data lost to endpoint failures. Not your fault, but it's what separates this from a 5.

Cite this work

@misc {

title={

(HckPrj) Correspondence Audits for Secretly Loyal Language Models: Blind, Base-Calibrated Detection Across Four Probe Families

},

author={

Jason Hung

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.