Containment Receipts: Offline-Verifiable Evidence for AI Agent Sandbox Compliance
Osman Yılmaz · Team CoRe
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Containment Receipts (CoRe) is a prototype for externally auditing whether an AI-agent evaluation remained within declared sandbox boundaries without trusting the agent transcript or accessing the evaluator’s internal network. CoRe binds policy, run bounds, observed events, and explicit telemetry-coverage claims into an Ed25519-signed receipt that can be verified offline. In a preregistered controlled Docker benchmark of 180 runs, a presence-only baseline incorrectly returned VERIFIED for 100 of 140 runs that should not have been verified (71.43%), while CoRe returned 0/140. Secondary studies evaluate replay, sequence anomalies, and coherent but false collector accounts using freshness and optional witness consistency. The result is a fail-closed, adoption-oriented verification profile for externally auditable containment evidence.
Reviews
The distinction between “no violation was observed” and “we have evidence that the entire run was observed” is useful and well operationalized. The explicit coverage checks are highly efficient, adding only fractions of a millisecond to the offline verification time.
The coverage-vs-presence distinction is compelling, but the experimental baseline seems too weak to establish how much CoRe itself contributes. The “presence-only” verifier treats the existence of events as evidence availability, so it predictably fails on scenarios constructed around observation gaps. A stronger comparison would be against a realistic audit pipeline that already tracks collector health, sequence numbers, heartbeats, or known logging gaps. If explicit coverage metadata still produces a meaningful improvement over that baseline, the result would say much more about CoRe’s practical value rather than demonstrating that explicit coverage beats an intentionally coverage-unaware verifier.
Read full reviewShow less
Quite similar to many other projects in this sprint. I think the emphasis on offline verifiability is a bit overstated - it's just a cryptographic signature on the verification output, nothing that intrinsticly ties it cryptographically to the compliance status to serve as a proof for example - you still need to trust the verifier, just not to query it online.
The writing was ok overall, but the inclusion of references to specific phases throughout the code and paper was distracting and unnecessary. It's also a tell-tale sign of AI writing.
I think the question this project asks is interesting: did we log an agent run well enough for an investigator to assess it? The project creates signed receipts that summarize what was logged and whether monitoring had gaps. The report also explains an important limitation clearly: a signed, internally consistent account can still be false.
This could be a promising tool for checking whether an evaluation’s logs are complete enough to judge if an agent followed the rules, rather than proving that an agent stayed contained. The next question is whether it remains useful when another team collects the evidence under less controlled conditions. I would test whether it catches gaps that ordinary logging misses, without making legitimate evaluations too difficult to assess. This would help establish whether the extra assurance justifies the effort of producing and checking these receipts.
Cite this project
@misc{ylmaz2026containment,
title = {{Containment Receipts: Offline-Verifiable Evidence for AI Agent Sandbox Compliance}},
author = {Osman Yılmaz},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/containment-receipts-offlineverifiable-evidence-for-ai-agent-sandbox-compliance-kahj}},
url = {https://apartresearch.com/sprints/projects/containment-receipts-offlineverifiable-evidence-for-ai-agent-sandbox-compliance-kahj}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …