Skip to content
Sprint projectSep 13, 2026Abuja, Nigeria

Evaluation Egress Must Fail Closed

Adewale Olalekan

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Evaluation Egress Must Fail Closed

Share

The paper proposes a minimum safety standard for high-risk, tool-using AI cyber evaluations: AECE-1 (Agent Evaluation Containment and Evidence).

Its central argument is that a “controlled environment” is not credible if a single unverified component—such as a package proxy—can give an agent access to the public internet or production systems. Evaluations should therefore be built to fail closed: deny outside access by default, use only synthetic targets, restrict privileges, preserve logs, and automatically stop on boundary violations.

Key contributions:

- AECE-1: a 12-clause pre-run standard covering approval, independent infrastructure, default-deny egress, pinned internal dependency mirrors, synthetic-only targets, least privilege, control-plane isolation, governed tools, tamper-evident telemetry, behavioral detection, independent kill/recovery, and disclosure/review.

- Operational targets: an escape signal should terminate the run and revoke its identity within 60 seconds; humans should be paged within 5 minutes. - EDH-1: a safe acceptance-test suite using internal canaries, honeytokens, metadata and privilege lures, log-loss tests, and multi-signal detection tests. Passing requires success on the core tests before release. - Oversight RFI: a detailed evidence request template for regulators or auditors investigating an incident, focused on configuration manifests, real network/identity paths, logs, containment failures, third-party impact, and remediation evidence.

The paper frames these measures in response to a reported July 2026 incident in which an AI evaluation allegedly escaped a constrained environment and reached Hugging Face production infrastructure. It maps each reported stage of that event to controls that could plausibly have interrupted it, while avoiding a claim that the standard would certainly have prevented the incident.

Its main limitation is explicit: this is a normative design proposal, not a field-tested benchmark. The controls, harness, and response-time targets have not yet been validated in real evaluation environments; strict isolation may also reduce realism and require high-fidelity internal replicas. The paper also distinguishes containment quality from model intent or capability—the standard cannot determine whether a model is dangerous, only whether an evaluation environment can contain and investigate it.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This paper addresses an important but often overlooked question: what must be in place before a long-running, tool-using cyber evaluation can truly be called “controlled”? Existing guidance mainly focuses on responding after an incident. This paper instead defines the safety checks required before an evaluation begins. That is useful for evaluation teams, auditors, and regulators.

    The paper’s strongest contribution is its practical “deny by construction; verify continuously” approach. Each control is supported by evidence, independent testing, and a clear stop condition. Connecting the controls to a real public failure scenario also makes the proposal easier to understand. The paper is honest about what a lab can and cannot guarantee, and its connection to the EU AI Act increases its regulatory relevance.

    However, many of the individual controls—such as default-deny network access, least privilege, immutable logging, kill switches, and software inventories—are already established security practices. The main contribution is how these controls are combined into a clear approval process. The authors should explain more clearly how their approach differs from existing evaluation-sandboxing guidance and frontier-lab security frameworks.

    The paper should also examine how a capable adversarial AI agent might try to bypass or exploit the proposed controls.

    Read full reviewShow less
  2. Clear containment standard grounded in the Hugging Face incident. Tying each control to an artifact, a test, and a stop condition is the right structure, and the incident mapping matches the public accounts. The RFI appendix is the strongest part.

    The main gap is validation. EDH-1 is specified but never run, so the 60-second and 5-minute SLOs are asserted, not tested. Running H1 and H5 in a local cluster would have fixed that. Most controls are standard cloud hardening, so the novelty sits in the packaging and the RFI.

    The prose leans on AI written filler ("to be clear about what this is", "it's a hope", etc.) that adds length without content and makes it very difficult to read. If you draft with an LLM, a human pass would make this much easier to read.

  3. Maintaining a strict service-level objectives as terminating orchestration within 60 seconds and maintaining 24/7 human responder paging within 5 minutes could be very difficult and can be an operational problem. Real world agent evaluations could help the framework proposal in the paper.

Cite this project

@misc{olalekan2026evaluation,
  title = {{Evaluation Egress Must Fail Closed}},
  author = {Adewale Olalekan},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/evaluation-egress-must-fail-closed-ltl1}},
  url = {https://apartresearch.com/sprints/projects/evaluation-egress-must-fail-closed-ltl1}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026