Egress Conformance: An Externally Verifiable Containment Property for Evaluation Sandboxes
Nafees · Team NafeesAI
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
The July 2026 sandbox escape went through the package-registry cache proxy that was the evaluation's one permitted network egress. We specify that sanctioned egress must be a content-addressed, replay-only mirror with no code-execution path, written as six black-box predicates and executed against two reference mirrors. Five of six discriminate; the sixth passes for both and is reported as a necessary property rather than counted as a test. Mapped against the thirteen trust boundaries in Hugging Face's forensic timeline, the property covers the first, which we argue is correct scope: it is the only crossing inside the lab's control and the only one with no external observer.
Reviews
The proposed containment property P2 restricts package installation to artifacts available in a pre-approved snapshot.
While this reduces the attack surface of the package-egress channel, it can also restrict legitimate agent capabilities when completing an evaluation requires installing legit dependencies.
Therefore, the approach appears most applicable to evaluations where dependencies can be fixed in advance, while its applicability to more open-ended agent evaluations is less clear.
I enjoyed reading the paper and I liked that it was clear about what could and couldn’t actually be verified. I also think making the failed conformance testing a main finding was a good choice, especially as a component returning an error doesn’t necessarily mean the security control worked, it could simply mean the route doesn’t exist.
One point I’d consider is the difference between verifying the artefact and verifying what is actually running in the lab. The paper says the deployment is attested rather than verified, but I think there are two separate problems here.
Cite this project
@misc{nafees2026egress,
title = {{Egress Conformance: An Externally Verifiable Containment Property for Evaluation Sandboxes}},
author = {Nafees},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/egress-conformance-an-externally-verifiable-containment-property-for-evaluation-sandboxes-giqn}},
url = {https://apartresearch.com/sprints/projects/egress-conformance-an-externally-verifiable-containment-property-for-evaluation-sandboxes-giqn}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …