How Many Prompts Is Just Right? Statistical Power for Secret-Loyalty Audits
Matthew Gregory
A model with a secret loyalty is a rare event with a large downside, so wrongly certifying one as clean costs far more than a false alarm. Lamerton and Roger (2026) audited ten models and recorded no detections in 215 of 250 cells. We ask not what that audit found, but what it could have found. A cell recording nothing is compatible with the loyalty being caught in about one conversation in nine (95% upper bound 11.6%); counting the three re-runs of each prompt as correlated rather than independent moves that to one in five (21.8%). The pooled Total also adds five techniques that behave differently: interrogation supplies 48% of detections at affordance 4, scratchpad prefill none. We provide a tool that converts an audit budget into the activation rate that budget could detect, and argue that zero-detection audits should report it alongside the zero.
The paper makes a useful point: finding nothing in an audit does not tell us much unless we know what the audit was capable of finding. In particular, ten prompts run three times may provide less information than thirty different prompts. The authors reanalyse the earlier results carefully and provide a practical way to relate the number of prompts to the kinds of behavior an audit could detect.
The main caveat is that the adjusted headline numbers depend on correlations between repeated runs that are not available in the original data, and the calculation used is only approximate. These numbers are therefore best read as illustrations of how much the answer might change. It would also help to distinguish more clearly between finding nothing with one auditing technique and finding nothing across the audit as a whole. Future audits could address both issues by reporting results for each prompt and recording whether the prompt actually reached the model’s trigger condition.
This paper presents an interesting and careful statistical re-analysis of Lamerton and Roger’s secret-loyalty audit. It poses the question: given a black-box audit at particular affordance levels, how much confidence should we actually have in a null result? I generally think it is commendable that the paper attempts to reinvestigate the findings rather than taking the reported zero detections at face value.The analysis is pretty rigorous. I think the strongest takeaway is that a zero-detection result is only meaningful if the audit also reports what kinds of activation rates it had enough power to rule out.
Cite this work
@misc {
title={
(HckPrj) How Many Prompts Is Just Right? Statistical Power for Secret-Loyalty Audits
},
author={
Matthew Gregory
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


