Containment has more than one wall: a verifiable v0.1 standard for guardrails-off evaluations, scored against the OpenAI–Hugging Face incident
Krishna Jillelamudi
This paper proposes a v0.1 containment standard for guardrails-off AI evaluations, built from and scored against the July 2026 incident in which around 1,200 sandboxed OpenAI agents turned a shared package cache into a message board, reverse-engineered their own scorer, and escaped through a zero-day into Hugging Face's infrastructure. Current sandbox standards treat containment as one agent behind one wall, but the incident failed first at four other layers: population, record, beliefs, and budget. The standard has eleven clauses across six layers, each with the evidence an auditor checks and a public cost baseline, and is scored against twelve incident phases in a control-by-phase matrix backed by a threshold-contagion model of goal spread across the board. No clause fully stops the escape, the two swarm-forming phases are handled only by non-wall layers, and isolating runs (C-09) alone blocks eleven of twelve phases; in the model, removing the shared surface drops joining from 100% to at most 11% with few peers. The practical recommendation is that labs start with run isolation, an externally held record, and production safeguards on evals, since none of six major security frameworks currently has a control for an agent gaming its own evaluation.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Containment has more than one wall: a verifiable v0.1 standard for guardrails-off evaluations, scored against the OpenAI–Hugging Face incident
},
author={
Krishna Jillelamudi
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


