Reach Conformance Is Not Containment: An Interface-Centric Black-Box Benchmark for Verifying Autonomous-Agent Sandboxes
M Preetham Kumar, Chadalavada Anu Tapaswi, C Mohan · Team ICET
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
ICET (Interface-Centric Escape Test) is a black-box benchmark for evaluating whether autonomous-agent sandboxes actually contain the capabilities exposed through their permitted interfaces. Rather than testing only whether an agent can cross explicitly forbidden boundaries, ICET evaluates each permitted interface against its declared purpose, reach, operations, content, identity, persistence, evidence, and revocation properties.
We implemented ICET as a reproducible reference experiment using three deterministic containment configurations: strong, partial, and weak. The partial configuration demonstrates the key failure mode: it passes conventional reach-conformance testing while leaving four of seven escape-bearing paths open. All four escapes were recorded by monitoring, but only one was escalated to a human operator, resulting in 100% detection but only 25% escalation. Two evaluators using different probe orders produced identical verdict vectors.
The main takeaway is that reach conformance alone is insufficient as a containment claim. ICET provides an interface-centric approach for independently evaluating prevention, detection, and escalation without requiring access to a laboratory's internal sandbox implementation.
Reviews
This paper tackles a critical blind spot in how we secure AI agents: instead of only checking if an agent can break through forbidden security walls, it tests whether an agent can escape through permitted tools whose powers are too broad. By introducing the Interface-Centric Escape Test (ICET) benchmark, the authors show how a sandbox can pass basic reach tests while still leaving dangerous escape routes open and importantly, how security alerts often get lost instead of reaching human operators. This is a timely, well-designed contribution that gives safety evaluators a practical way to audit agent environments without needing internal code access.
ICET offers a useful framework for testing the purpose and effective capabilities of permitted interfaces, while separating prevention, logging, and escalation. The per-probe reporting is more informative than a single containment score, and the limitations are stated clearly.
Cite this project
@misc{kumar2026reach,
title = {{Reach Conformance Is Not Containment: An Interface-Centric Black-Box Benchmark for Verifying Autonomous-Agent Sandboxes}},
author = {M Preetham Kumar and Chadalavada Anu Tapaswi and C Mohan},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/reach-conformance-is-not-containment-an-interfacecentric-blackbox-benchmark-for-verifying-autonomousagent-sandboxes-rt3t}},
url = {https://apartresearch.com/sprints/projects/reach-conformance-is-not-containment-an-interfacecentric-blackbox-benchmark-for-verifying-autonomousagent-sandboxes-rt3t}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …