Skip to content
Sprint projectSep 12, 2026Aydın,Türkiye

Containment Receipts: Offline-Verifiable Evidence for AI Agent Sandbox Compliance

Osman Yılmaz · Team CoRe

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Containment Receipts: Offline-Verifiable Evidence for AI Agent Sandbox Compliance

Code (opens in new tab)
Share

Containment Receipts (CoRe) is a prototype for externally auditing whether an AI-agent evaluation remained within declared sandbox boundaries without trusting the agent transcript or accessing the evaluator’s internal network. CoRe binds policy, run bounds, observed events, and explicit telemetry-coverage claims into an Ed25519-signed receipt that can be verified offline. In a preregistered controlled Docker benchmark of 180 runs, a presence-only baseline incorrectly returned VERIFIED for 100 of 140 runs that should not have been verified (71.43%), while CoRe returned 0/140. Secondary studies evaluate replay, sequence anomalies, and coherent but false collector accounts using freshness and optional witness consistency. The result is a fail-closed, adoption-oriented verification profile for externally auditable containment evidence.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The distinction between “no violation was observed” and “we have evidence that the entire run was observed” is useful and well operationalized. The explicit coverage checks are highly efficient, adding only fractions of a millisecond to the offline verification time.

    The coverage-vs-presence distinction is compelling, but the experimental baseline seems too weak to establish how much CoRe itself contributes. The “presence-only” verifier treats the existence of events as evidence availability, so it predictably fails on scenarios constructed around observation gaps. A stronger comparison would be against a realistic audit pipeline that already tracks collector health, sequence numbers, heartbeats, or known logging gaps. If explicit coverage metadata still produces a meaningful improvement over that baseline, the result would say much more about CoRe’s practical value rather than demonstrating that explicit coverage beats an intentionally coverage-unaware verifier.

    Read full reviewShow less
  2. Quite similar to many other projects in this sprint. I think the emphasis on offline verifiability is a bit overstated - it's just a cryptographic signature on the verification output, nothing that intrinsticly ties it cryptographically to the compliance status to serve as a proof for example - you still need to trust the verifier, just not to query it online.

    The writing was ok overall, but the inclusion of references to specific phases throughout the code and paper was distracting and unnecessary. It's also a tell-tale sign of AI writing.

  3. I think the question this project asks is interesting: did we log an agent run well enough for an investigator to assess it? The project creates signed receipts that summarize what was logged and whether monitoring had gaps. The report also explains an important limitation clearly: a signed, internally consistent account can still be false.

    This could be a promising tool for checking whether an evaluation’s logs are complete enough to judge if an agent followed the rules, rather than proving that an agent stayed contained. The next question is whether it remains useful when another team collects the evidence under less controlled conditions. I would test whether it catches gaps that ordinary logging misses, without making legitimate evaluations too difficult to assess. This would help establish whether the extra assurance justifies the effort of producing and checking these receipts.

Cite this project

@misc{ylmaz2026containment,
  title = {{Containment Receipts: Offline-Verifiable Evidence for AI Agent Sandbox Compliance}},
  author = {Osman Yılmaz},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/containment-receipts-offlineverifiable-evidence-for-ai-agent-sandbox-compliance-kahj}},
  url = {https://apartresearch.com/sprints/projects/containment-receipts-offlineverifiable-evidence-for-ai-agent-sandbox-compliance-kahj}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026