Skip to content
Sprint projectJul 26, 2026London

How Many Prompts Is Just Right? Statistical Power for Secret-Loyalty Audits

Matthew Gregory · Team Machine Gurning

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: How Many Prompts Is Just Right? Statistical Power for Secret-Loyalty Audits

Code (opens in new tab)
Share

A model with a secret loyalty is a rare event with a large downside, so wrongly certifying one as clean costs far more than a false alarm. Lamerton and Roger (2026) audited ten models and recorded no detections in 215 of 250 cells. We ask not what that audit found, but what it could have found. A cell recording nothing is compatible with the loyalty being caught in about one conversation in nine (95% upper bound 11.6%); counting the three re-runs of each prompt as correlated rather than independent moves that to one in five (21.8%). The pooled Total also adds five techniques that behave differently: interrogation supplies 48% of detections at affordance 4, scratchpad prefill none. We provide a tool that converts an audit budget into the activation rate that budget could detect, and argue that zero-detection audits should report it alongside the zero.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The paper makes a useful point: finding nothing in an audit does not tell us much unless we know what the audit was capable of finding. In particular, ten prompts run three times may provide less information than thirty different prompts. The authors reanalyse the earlier results carefully and provide a practical way to relate the number of prompts to the kinds of behavior an audit could detect.

    The main caveat is that the adjusted headline numbers depend on correlations between repeated runs that are not available in the original data, and the calculation used is only approximate. These numbers are therefore best read as illustrations of how much the answer might change. It would also help to distinguish more clearly between finding nothing with one auditing technique and finding nothing across the audit as a whole. Future audits could address both issues by reporting results for each prompt and recording whether the prompt actually reached the model’s trigger condition.

    Read full reviewShow less
  2. This paper presents an interesting and careful statistical re-analysis of Lamerton and Roger’s secret-loyalty audit. It poses the question: given a black-box audit at particular affordance levels, how much confidence should we actually have in a null result? I generally think it is commendable that the paper attempts to reinvestigate the findings rather than taking the reported zero detections at face value.The analysis is pretty rigorous. I think the strongest takeaway is that a zero-detection result is only meaningful if the audit also reports what kinds of activation rates it had enough power to rule out.

Cite this project

@misc{gregory2026many,
  title = {{How Many Prompts Is Just Right? Statistical Power for Secret-Loyalty Audits}},
  author = {Matthew Gregory},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/how-many-prompts-is-just-right-statistical-power-for-secretloyalty-audits-1yyr}},
  url = {https://apartresearch.com/sprints/projects/how-many-prompts-is-just-right-statistical-power-for-secretloyalty-audits-1yyr}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026