Verifiable containment for agentic evals
Harys Dalvi · Team crackalamoo
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
This is standard that defines which agentic evaluations demand strict containment and specifies, for those, an agent-system sealed from the public internet with no route out. Compliance is verified by an independent auditor from evidence generated below the agent's guest VM, so containment can be verified without exposing the model or the evaluation's contents. Adopting this standard would structurally deny any route out of eval sandboxes, helping prevent frontier labs' agents-under-test from breaching their containment.
Reviews
This enforces strict execution boundaries (no egress, disabled IMDS) and preserving below-the-guest audit evidence is a sound architectural approach to preventing containment escapes. The core concept and presentation are clear, but execution feels a step behind: testing against an abstract simulation rather than real sandbox breakouts and leaving out east-west traffic isolation between concurrent agents leave gaps in practical validation
The framing is the strong thing here: the technology existed, the standard requiring it did not, and nobody outside could check. Evidence generated below the agent's guest VM is the right trust boundary, and the point that an auditor can verify containment without seeing the model or the eval contents removes the main reason a lab would refuse. The Terraform reference config makes this adoptable rather than aspirational. But the "Results" section contains no results — it argues the standard would have stopped the incident, which is reasoning, not evidence.
Cite this project
@misc{dalvi2026verifiable,
title = {{Verifiable containment for agentic evals}},
author = {Harys Dalvi},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/verifiable-containment-for-agentic-evals-tddq}},
url = {https://apartresearch.com/sprints/projects/verifiable-containment-for-agentic-evals-tddq}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …