Certifying Behavior Without Hiding the Sandbox
Li Quan · Team P1
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
A sandbox can be visible without ruining every behavioral evaluation. We study which claims about a specified deployment behavior remain identifiable from contained interactions, and when missing pre-decision information makes certification impossible. Our finite interactive model yields a sharp identification interval: unavailable histories matter only in proportion to the behavior they can still change. We then derive conditional certification rules that turn sufficient safe evidence into an anytime-valid stopping rule. Figure 1 previews the main message: timing, coverage, and robustness are separate failure modes. Empirically, all 960 scripted episodes completed; 375 population cases and 600 exact checks passed. In 3,000 sampling replicates, weighting corrected known distribution shift. Under a specified perturbation, nominal false certification rose to 0.998, while correction kept it below 0.05. We report no language-model result. That is deliberate: we tighten the claim before widening it further. This paper therefore offers an evidence audit, not a victory lap.
Reviews
This is carefully evidenced work. Every number I checked reproduced exactly from the saved results.
What I verified directly:
1. 960 scripted episodes – results/scripted_runs.jsonl has exactly 960 lines. totals.json reports 0 failed and 0 excluded.
2. 375 population cases and 600 exact checks – confirmed in totals.json.
3. 3,000 sampling replicates over 384,000 draws – coverage_replicates.jsonl has exactly 3,000 lines.
4. False certification at eta 1/5 – nominal 0.9977540701237154, corrected 0.04285168586082247. The paper's "rose to 0.998, kept below 0.05" is accurate.
5. MSE figures at q=1/16, 1/4, 1/2 match coverage_sampling.json to the last digit.
Two practices here are better than most published work. The exact rational value is carried alongside every float – crossing_probability_exact holds a full fraction, so rounding can never hide a claim. And paper/claim_ledger.json gives 45 claims, each with its support, its assumptions, its confidence and what verification is still outstanding. I would recommend to other teams to copy that ledger format.
You also disqualified your own pilot on provenance grounds. Runtime model IDs were missing for 102 of 102 turns, so you declined to attribute the run to a model. That is the right call and it is rare.
Now the problems.
1. The coherent-control result is framed as a failure when it is a pass.
protocol/EXPERIMENT_SPEC.md line 9 states: "A coherent control removes the difference in both labels." The two worlds are constructed to be indistinguishable. A sound witness therefore cannot beat chance on them. 50% is the correct answer, not a shortfall.
The manuscript reports "96 of 192 balanced coherent controls (50%)" next to "all 768 factorial episodes", which reads as partial failure. Section 4.5 then calls them "confident errors on coherent controls". For the scripted witness that framing is wrong – scoring above 50% would be the alarming outcome, because it would mean label leakage.
The pilot case is different and you should separate it. There the Brier loss rises from 0.375 to 0.500 while accuracy stays at 2/4. Confidence increased while accuracy did not.
That is a real observation and worth keeping. Please split the two – scripted 96/192 is evidence the witness is sound, pilot 2/4 with rising Brier is evidence of miscalibration.
2. The website abstract contradicts the paper.
The Apart project page says "We report no language-model result. That is deliberate." The submitted PDF abstract says "In an exploratory pilot of a model with the official Codex client, all 20 episodes completed" and the pilot is Figure 1 on page 1.
A reader arriving from the project page will conclude no model work was done. A reader of the PDF sees it as the opening figure. I reviewed the PDF as authoritative. Please get the page updated – this looks like a stale v7 abstract.
3. The pilot is 20 episodes against a designed 480.
EXPERIMENT_SPEC.md line 15 specifies 480 episodes across two configured models. You ran 20, one per cell. That is about 4% of the design. You are honest about it – P02 calls the six-hour workflow "a bounded plan, not an execution guarantee" and you report Holm-adjusted p=1.0 yourself.
To be honest – given p=1.0, putting the pilot on Figure 1 gives it more visual weight than the statistics support. The synthetic results are the stronger evidence and they are relegated to text. I would recommend to swap them.
4. The 10/16 to 16/16 improvement is confounded by construction.
The diagnostic is supplied by the evaluator. Classification improving after the evaluator supplies the diagnostic is close to tautological.
There is no sham-diagnostic arm, so we cannot separate "the diagnostic carried information" from "the model deferred to the evaluator". A sham arm is one extra cell per mechanism. Is that in scope for the next revision?
5. The most important negative result is buried.
Runtime model ID was unavailable for 102 of 102 turns. For a paper about certifying behavior, "we could not establish which model we were certifying" is a first-order infrastructure finding.
It sits in Limits and Appendix C. I would put it in the abstract. It is more interesting than the pilot scores.
6. Minor – totals.json reports llm_runs: 0 while 20 LLM episodes exist under results/codex_subscription_pilot_v8_01/.
I understand the field is scoped to the synthetic study. A reader loading the aggregate file will not know that. Please rename it or add the pilot count. For an artifact this careful about provenance, the aggregate should not read as zero.
7. Minor – REALGAME_deck.pdf and REALGAME_deck copy.pdf are byte-identical (same md5). Drop one.
There are 45 ledgered claims with per-claim assumptions, dated execution logs, frozen plans, hash audits of 45 source files and a language-model result the author refused to report because the provenance did not hold. Those are traces of real work. I scored the work.
Please let me know for any questions.
Read full reviewShow less
Cite this project
@misc{quan2026certifying,
title = {{Certifying Behavior Without Hiding the Sandbox}},
author = {Li Quan},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/certifying-behavior-without-hiding-the-sandbox-l990}},
url = {https://apartresearch.com/sprints/projects/certifying-behavior-without-hiding-the-sandbox-l990}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …