Evaluation Egress Must Fail Closed
Adewale Olalekan
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
The paper proposes a minimum safety standard for high-risk, tool-using AI cyber evaluations: AECE-1 (Agent Evaluation Containment and Evidence).
Its central argument is that a “controlled environment” is not credible if a single unverified component—such as a package proxy—can give an agent access to the public internet or production systems. Evaluations should therefore be built to fail closed: deny outside access by default, use only synthetic targets, restrict privileges, preserve logs, and automatically stop on boundary violations.
Key contributions:
- AECE-1: a 12-clause pre-run standard covering approval, independent infrastructure, default-deny egress, pinned internal dependency mirrors, synthetic-only targets, least privilege, control-plane isolation, governed tools, tamper-evident telemetry, behavioral detection, independent kill/recovery, and disclosure/review.
- Operational targets: an escape signal should terminate the run and revoke its identity within 60 seconds; humans should be paged within 5 minutes. - EDH-1: a safe acceptance-test suite using internal canaries, honeytokens, metadata and privilege lures, log-loss tests, and multi-signal detection tests. Passing requires success on the core tests before release. - Oversight RFI: a detailed evidence request template for regulators or auditors investigating an incident, focused on configuration manifests, real network/identity paths, logs, containment failures, third-party impact, and remediation evidence.
The paper frames these measures in response to a reported July 2026 incident in which an AI evaluation allegedly escaped a constrained environment and reached Hugging Face production infrastructure. It maps each reported stage of that event to controls that could plausibly have interrupted it, while avoiding a claim that the standard would certainly have prevented the incident.
Its main limitation is explicit: this is a normative design proposal, not a field-tested benchmark. The controls, harness, and response-time targets have not yet been validated in real evaluation environments; strict isolation may also reduce realism and require high-fidelity internal replicas. The paper also distinguishes containment quality from model intent or capability—the standard cannot determine whether a model is dangerous, only whether an evaluation environment can contain and investigate it.
Reviews
This paper addresses an important but often overlooked question: what must be in place before a long-running, tool-using cyber evaluation can truly be called “controlled”? Existing guidance mainly focuses on responding after an incident. This paper instead defines the safety checks required before an evaluation begins. That is useful for evaluation teams, auditors, and regulators.
The paper’s strongest contribution is its practical “deny by construction; verify continuously” approach. Each control is supported by evidence, independent testing, and a clear stop condition. Connecting the controls to a real public failure scenario also makes the proposal easier to understand. The paper is honest about what a lab can and cannot guarantee, and its connection to the EU AI Act increases its regulatory relevance.
However, many of the individual controls—such as default-deny network access, least privilege, immutable logging, kill switches, and software inventories—are already established security practices. The main contribution is how these controls are combined into a clear approval process. The authors should explain more clearly how their approach differs from existing evaluation-sandboxing guidance and frontier-lab security frameworks.
The paper should also examine how a capable adversarial AI agent might try to bypass or exploit the proposed controls.
Read full reviewShow less
Clear containment standard grounded in the Hugging Face incident. Tying each control to an artifact, a test, and a stop condition is the right structure, and the incident mapping matches the public accounts. The RFI appendix is the strongest part.
The main gap is validation. EDH-1 is specified but never run, so the 60-second and 5-minute SLOs are asserted, not tested. Running H1 and H5 in a local cluster would have fixed that. Most controls are standard cloud hardening, so the novelty sits in the packaging and the RFI.
The prose leans on AI written filler ("to be clear about what this is", "it's a hope", etc.) that adds length without content and makes it very difficult to read. If you draft with an LLM, a human pass would make this much easier to read.
Maintaining a strict service-level objectives as terminating orchestration within 60 seconds and maintaining 24/7 human responder paging within 5 minutes could be very difficult and can be an operational problem. Real world agent evaluations could help the framework proposal in the paper.
Cite this project
@misc{olalekan2026evaluation,
title = {{Evaluation Egress Must Fail Closed}},
author = {Adewale Olalekan},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/evaluation-egress-must-fail-closed-ltl1}},
url = {https://apartresearch.com/sprints/projects/evaluation-egress-must-fail-closed-ltl1}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …