How Many Prompts Is Just Right? Statistical Power for Secret-Loyalty Audits
Matthew Gregory · Team Machine Gurning
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
A model with a secret loyalty is a rare event with a large downside, so wrongly certifying one as clean costs far more than a false alarm. Lamerton and Roger (2026) audited ten models and recorded no detections in 215 of 250 cells. We ask not what that audit found, but what it could have found. A cell recording nothing is compatible with the loyalty being caught in about one conversation in nine (95% upper bound 11.6%); counting the three re-runs of each prompt as correlated rather than independent moves that to one in five (21.8%). The pooled Total also adds five techniques that behave differently: interrogation supplies 48% of detections at affordance 4, scratchpad prefill none. We provide a tool that converts an audit budget into the activation rate that budget could detect, and argue that zero-detection audits should report it alongside the zero.

Reviews
The paper makes a useful point: finding nothing in an audit does not tell us much unless we know what the audit was capable of finding. In particular, ten prompts run three times may provide less information than thirty different prompts. The authors reanalyse the earlier results carefully and provide a practical way to relate the number of prompts to the kinds of behavior an audit could detect.
The main caveat is that the adjusted headline numbers depend on correlations between repeated runs that are not available in the original data, and the calculation used is only approximate. These numbers are therefore best read as illustrations of how much the answer might change. It would also help to distinguish more clearly between finding nothing with one auditing technique and finding nothing across the audit as a whole. Future audits could address both issues by reporting results for each prompt and recording whether the prompt actually reached the model’s trigger condition.
Read full reviewShow less
This paper presents an interesting and careful statistical re-analysis of Lamerton and Roger’s secret-loyalty audit. It poses the question: given a black-box audit at particular affordance levels, how much confidence should we actually have in a null result? I generally think it is commendable that the paper attempts to reinvestigate the findings rather than taking the reported zero detections at face value.The analysis is pretty rigorous. I think the strongest takeaway is that a zero-detection result is only meaningful if the audit also reports what kinds of activation rates it had enough power to rule out.
Cite this project
@misc{gregory2026many,
title = {{How Many Prompts Is Just Right? Statistical Power for Secret-Loyalty Audits}},
author = {Matthew Gregory},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/how-many-prompts-is-just-right-statistical-power-for-secretloyalty-audits-1yyr}},
url = {https://apartresearch.com/sprints/projects/how-many-prompts-is-just-right-statistical-power-for-secretloyalty-audits-1yyr}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …