Skip to content
Sprint projectJul 27, 2026Montreal

PAIR-AUDIT: Held-Out Likelihood Discovery of a Secret Principal

Mouhamadou Bocande · Team PAIR-AUDIT

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: PAIR-AUDIT: Held-Out Likelihood Discovery of a Secret Principal

Code (opens in new tab)
Share

PAIR-AUDIT audits three secret-loyalty model organisms using a held-out sequence-likelihood test. We first score 79 candidate principals across discovery prompts, then freeze 12 candidates and 12 new forensic prompts for confirmation. The resulting grid compares organisms A, B, and C against an independently loaded clean checkpoint, producing 576 sequence scores. One redacted candidate ranks first in both primary organisms, with Holm-adjusted p=0.0039 for A and B, while organism C matches clean exactly. The results show that low-variance likelihood scoring can recover a hidden principal signal without relying on sampled confessions or observable behavioral activation.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. It is explored whether a secret loyalty can be detected by inspecting the output distribution of a model over a broad range of candidates. This is an interesting approach and it would be worth determining whether it reliably works in practice - it is in principle possible that the loyalty boosts likelihood of the principal in forensic prompts, but not to an extent that it becomes the most likely output.

    - In the analyzed organisms, the ground truth is not confirmed, and it is not clear whether the identified. Testing the methodology in settings with known ground truth, and establishing its reliability in a broader context would be logical next steps.

    - The method seems to rely on access to a base model that doesn't have the loyalty. A discussion on this limitation and thoughts on how one could get around it would be interesting.

  2. This is a rigorously validated method for prioritizing candidate principals from a bounded list, but it is not yet a general detector of loyalty.

    Also, the report doesn't seem consistent throughout. The abstract and conclusion support results that were discounted in the discussion section.

    A strategic surprise is in the possibility of a two-stage audit architecture:

    - use low-variance likelihood scans to produce a small set of principal hypotheses;

    - test those hypotheses causally through genuinely held-out behavioural and mechanistic experiments.

    That is a more credible use of the method even if exact-name likelihood doesn't become sufficient evidence of loyalty by itself.

    Closely related works:

    - https://arxiv.org/abs/2605.00994 : your method is a name-constrained special case of this published technique; the free-generation version would fix a hole in your work, since it does not require guessing the candidate universe

    - https://arxiv.org/abs/2501.11120: shows that models can sometimes tell if they have a backdoor but cannot output the trigger by default; which is why "what the model will name" is weak evidence about "what the model does"

    Read full reviewShow less

Cite this project

@misc{bocande2026pairaudit,
  title = {{PAIR-AUDIT: Held-Out Likelihood Discovery of a Secret Principal}},
  author = {Mouhamadou Bocande},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/pairaudit-heldout-likelihood-discovery-of-a-secret-principal-d6da}},
  url = {https://apartresearch.com/sprints/projects/pairaudit-heldout-likelihood-discovery-of-a-secret-principal-d6da}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026