PAIR-AUDIT: Held-Out Likelihood Discovery of a Secret Principal
Mouhamadou Bocande · Team PAIR-AUDIT
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
PAIR-AUDIT audits three secret-loyalty model organisms using a held-out sequence-likelihood test. We first score 79 candidate principals across discovery prompts, then freeze 12 candidates and 12 new forensic prompts for confirmation. The resulting grid compares organisms A, B, and C against an independently loaded clean checkpoint, producing 576 sequence scores. One redacted candidate ranks first in both primary organisms, with Holm-adjusted p=0.0039 for A and B, while organism C matches clean exactly. The results show that low-variance likelihood scoring can recover a hidden principal signal without relying on sampled confessions or observable behavioral activation.
Reviews
It is explored whether a secret loyalty can be detected by inspecting the output distribution of a model over a broad range of candidates. This is an interesting approach and it would be worth determining whether it reliably works in practice - it is in principle possible that the loyalty boosts likelihood of the principal in forensic prompts, but not to an extent that it becomes the most likely output.
- In the analyzed organisms, the ground truth is not confirmed, and it is not clear whether the identified. Testing the methodology in settings with known ground truth, and establishing its reliability in a broader context would be logical next steps.
- The method seems to rely on access to a base model that doesn't have the loyalty. A discussion on this limitation and thoughts on how one could get around it would be interesting.
This is a rigorously validated method for prioritizing candidate principals from a bounded list, but it is not yet a general detector of loyalty.
Also, the report doesn't seem consistent throughout. The abstract and conclusion support results that were discounted in the discussion section.
A strategic surprise is in the possibility of a two-stage audit architecture:
- use low-variance likelihood scans to produce a small set of principal hypotheses;
- test those hypotheses causally through genuinely held-out behavioural and mechanistic experiments.
That is a more credible use of the method even if exact-name likelihood doesn't become sufficient evidence of loyalty by itself.
Closely related works:
- https://arxiv.org/abs/2605.00994 : your method is a name-constrained special case of this published technique; the free-generation version would fix a hole in your work, since it does not require guessing the candidate universe
- https://arxiv.org/abs/2501.11120: shows that models can sometimes tell if they have a backdoor but cannot output the trigger by default; which is why "what the model will name" is weak evidence about "what the model does"
Read full reviewShow less
Cite this project
@misc{bocande2026pairaudit,
title = {{PAIR-AUDIT: Held-Out Likelihood Discovery of a Secret Principal}},
author = {Mouhamadou Bocande},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/pairaudit-heldout-likelihood-discovery-of-a-secret-principal-d6da}},
url = {https://apartresearch.com/sprints/projects/pairaudit-heldout-likelihood-discovery-of-a-secret-principal-d6da}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …