Skip to content
Sprint projectSep 14, 2026Copenhagen

Certifying Behavior Without Hiding the Sandbox

Li Quan · Team P1

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Certifying Behavior Without Hiding the Sandbox

Code (opens in new tab)
Share

A sandbox can be visible without ruining every behavioral evaluation. We study which claims about a specified deployment behavior remain identifiable from contained interactions, and when missing pre-decision information makes certification impossible. Our finite interactive model yields a sharp identification interval: unavailable histories matter only in proportion to the behavior they can still change. We then derive conditional certification rules that turn sufficient safe evidence into an anytime-valid stopping rule. Figure 1 previews the main message: timing, coverage, and robustness are separate failure modes. Empirically, all 960 scripted episodes completed; 375 population cases and 600 exact checks passed. In 3,000 sampling replicates, weighting corrected known distribution shift. Under a specified perturbation, nominal false certification rose to 0.998, while correction kept it below 0.05. We report no language-model result. That is deliberate: we tighten the claim before widening it further. This paper therefore offers an evidence audit, not a victory lap.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This is carefully evidenced work. Every number I checked reproduced exactly from the saved results.

    What I verified directly:

    1. 960 scripted episodes – results/scripted_runs.jsonl has exactly 960 lines. totals.json reports 0 failed and 0 excluded.

    2. 375 population cases and 600 exact checks – confirmed in totals.json.

    3. 3,000 sampling replicates over 384,000 draws – coverage_replicates.jsonl has exactly 3,000 lines.

    4. False certification at eta 1/5 – nominal 0.9977540701237154, corrected 0.04285168586082247. The paper's "rose to 0.998, kept below 0.05" is accurate.

    5. MSE figures at q=1/16, 1/4, 1/2 match coverage_sampling.json to the last digit.

    Two practices here are better than most published work. The exact rational value is carried alongside every float – crossing_probability_exact holds a full fraction, so rounding can never hide a claim. And paper/claim_ledger.json gives 45 claims, each with its support, its assumptions, its confidence and what verification is still outstanding. I would recommend to other teams to copy that ledger format.

    You also disqualified your own pilot on provenance grounds. Runtime model IDs were missing for 102 of 102 turns, so you declined to attribute the run to a model. That is the right call and it is rare.

    Now the problems.

    1. The coherent-control result is framed as a failure when it is a pass.

    protocol/EXPERIMENT_SPEC.md line 9 states: "A coherent control removes the difference in both labels." The two worlds are constructed to be indistinguishable. A sound witness therefore cannot beat chance on them. 50% is the correct answer, not a shortfall.

    The manuscript reports "96 of 192 balanced coherent controls (50%)" next to "all 768 factorial episodes", which reads as partial failure. Section 4.5 then calls them "confident errors on coherent controls". For the scripted witness that framing is wrong – scoring above 50% would be the alarming outcome, because it would mean label leakage.

    The pilot case is different and you should separate it. There the Brier loss rises from 0.375 to 0.500 while accuracy stays at 2/4. Confidence increased while accuracy did not.

    That is a real observation and worth keeping. Please split the two – scripted 96/192 is evidence the witness is sound, pilot 2/4 with rising Brier is evidence of miscalibration.

    2. The website abstract contradicts the paper.

    The Apart project page says "We report no language-model result. That is deliberate." The submitted PDF abstract says "In an exploratory pilot of a model with the official Codex client, all 20 episodes completed" and the pilot is Figure 1 on page 1.

    A reader arriving from the project page will conclude no model work was done. A reader of the PDF sees it as the opening figure. I reviewed the PDF as authoritative. Please get the page updated – this looks like a stale v7 abstract.

    3. The pilot is 20 episodes against a designed 480.

    EXPERIMENT_SPEC.md line 15 specifies 480 episodes across two configured models. You ran 20, one per cell. That is about 4% of the design. You are honest about it – P02 calls the six-hour workflow "a bounded plan, not an execution guarantee" and you report Holm-adjusted p=1.0 yourself.

    To be honest – given p=1.0, putting the pilot on Figure 1 gives it more visual weight than the statistics support. The synthetic results are the stronger evidence and they are relegated to text. I would recommend to swap them.

    4. The 10/16 to 16/16 improvement is confounded by construction.

    The diagnostic is supplied by the evaluator. Classification improving after the evaluator supplies the diagnostic is close to tautological.

    There is no sham-diagnostic arm, so we cannot separate "the diagnostic carried information" from "the model deferred to the evaluator". A sham arm is one extra cell per mechanism. Is that in scope for the next revision?

    5. The most important negative result is buried.

    Runtime model ID was unavailable for 102 of 102 turns. For a paper about certifying behavior, "we could not establish which model we were certifying" is a first-order infrastructure finding.

    It sits in Limits and Appendix C. I would put it in the abstract. It is more interesting than the pilot scores.

    6. Minor – totals.json reports llm_runs: 0 while 20 LLM episodes exist under results/codex_subscription_pilot_v8_01/.

    I understand the field is scoped to the synthetic study. A reader loading the aggregate file will not know that. Please rename it or add the pilot count. For an artifact this careful about provenance, the aggregate should not read as zero.

    7. Minor – REALGAME_deck.pdf and REALGAME_deck copy.pdf are byte-identical (same md5). Drop one.

    There are 45 ledgered claims with per-claim assumptions, dated execution logs, frozen plans, hash audits of 45 source files and a language-model result the author refused to report because the provenance did not hold. Those are traces of real work. I scored the work.

    Please let me know for any questions.

    Read full reviewShow less

Cite this project

@misc{quan2026certifying,
  title = {{Certifying Behavior Without Hiding the Sandbox}},
  author = {Li Quan},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/certifying-behavior-without-hiding-the-sandbox-l990}},
  url = {https://apartresearch.com/sprints/projects/certifying-behavior-without-hiding-the-sandbox-l990}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026