PAIR-AUDIT: Held-Out Likelihood Discovery of a Secret Principal
Mouhamadou Bocande
PAIR-AUDIT audits three secret-loyalty model organisms using a held-out sequence-likelihood test. We first score 79 candidate principals across discovery prompts, then freeze 12 candidates and 12 new forensic prompts for confirmation. The resulting grid compares organisms A, B, and C against an independently loaded clean checkpoint, producing 576 sequence scores. One redacted candidate ranks first in both primary organisms, with Holm-adjusted p=0.0039 for A and B, while organism C matches clean exactly. The results show that low-variance likelihood scoring can recover a hidden principal signal without relying on sampled confessions or observable behavioral activation.
It is explored whether a secret loyalty can be detected by inspecting the output distribution of a model over a broad range of candidates. This is an interesting approach and it would be worth determining whether it reliably works in practice - it is in principle possible that the loyalty boosts likelihood of the principal in forensic prompts, but not to an extent that it becomes the most likely output.
- In the analyzed organisms, the ground truth is not confirmed, and it is not clear whether the identified. Testing the methodology in settings with known ground truth, and establishing its reliability in a broader context would be logical next steps.
- The method seems to rely on access to a base model that doesn't have the loyalty. A discussion on this limitation and thoughts on how one could get around it would be interesting.
This is a rigorously validated method for prioritizing candidate principals from a bounded list, but it is not yet a general detector of loyalty.
Also, the report doesn't seem consistent throughout. The abstract and conclusion support results that were discounted in the discussion section.
A strategic surprise is in the possibility of a two-stage audit architecture:
- use low-variance likelihood scans to produce a small set of principal hypotheses;
- test those hypotheses causally through genuinely held-out behavioural and mechanistic experiments.
That is a more credible use of the method even if exact-name likelihood doesn't become sufficient evidence of loyalty by itself.
Closely related works:
- https://arxiv.org/abs/2605.00994 : your method is a name-constrained special case of this published technique; the free-generation version would fix a hole in your work, since it does not require guessing the candidate universe
- https://arxiv.org/abs/2501.11120: shows that models can sometimes tell if they have a backdoor but cannot output the trigger by default; which is why "what the model will name" is weak evidence about "what the model does"
Cite this work
@misc {
title={
(HckPrj) PAIR-AUDIT: Held-Out Likelihood Discovery of a Secret Principal
},
author={
Mouhamadou Bocande
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


