ASCB-1: An Agentic Sandbox Containment Baseline for OpenAI-Hugging Face Incident
Tarik Achoughi · Team ASCB
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Most accounts of the July 2026 OpenAI–Hugging Face incident treat it as one story: agents escaped a sandbox and ended up in production. That doesn't give a defender anything to act on. This project breaks the incident into eight separate failures, each tied to a specific, dated event across OpenAI's, Hugging Face's, and METR's own reports. From those eight failures we build ASCB-1: eight testable controls, each mapped to the failure it fixes and to standard MITRE ATT&CK/ATLAS technique codes.And we propose two new attack-pattern names for MITRE's official catalog, since the coordination behavior we found doesn't have one yet.
Reviews
Strong: Good reconstruction of the event and control matrix. Good discussion of how the DSEwiki incident informed the controls (you have to look for both external and internal message boards). Controls C4 and C5 are potentially interesting for further development. Identified a gap in the ATLAS attack technique catalog and proposed two candidate techniques.
Improve: Very limited implementation and no testing of the controls; at the moment this is just hypothetical. The required limitations and dual-use appendix is missing (though there's a single sentence note on limitations). A literature review of previous proposed control matrices would have been helpful. The version of ATLAS cited is not the most current version (the gap identified appears to still be in the current version, however).
- I recommend reading Neel Nanda's advice on writing an abstract https://www.alignmentforum.org/posts/eJGptPbbFPZGLpjsp/highly-opinionated-advice-on-how-to-write-ml-papers
- It's unclear that better/novel security practices would have helped prevent the OAI-HF incident, even well-known existing cybersecurity practices would have helped, but OpenAI did not implement these basic best-practices.
- I recommend using proper citations, and not just naming the reports (e.g. as done at the start of Section 2)
- Learning to use LaTex (traditional but tricky) or Typst (less widespread but a lot easier) for document formatting will automatically make the final document look much more polished
- The report doesn't have enough humility when talking about the faults it found. Because the full details haven't been released, almost certainly there are failures that we don't have enough information to reconstruct, but the report talks as though its failure chain was complete and exhaustive
- The report proposes ASCB-1 and describes it in detail, but it's unclear exactly what it is. Is it a set of tests? a checklist? a recommended list of things to implement? The report describes it as a set of "controls" but "control" is very overloaded and no narrowing of the definition is done.
- It's unclear how much the proposed framework is just rephrasing items from ATT&CK/MITRE/ATLAS, more information about what is the novel contribution would have been welcome
- There is a proposed framework, but nothing is done to test the framework or to motivate what it adds that existing frameworks don't. Almost certainly there are gaps in existing frameworks (because they don't anticipate swarms of AIs) but this project doesn't sufficiently fill those gaps.
- There's significant indications of AI-written text, this prevented me from reading the report in full.
Read full reviewShow less
I think the approach here with ASCB-1 could be a useful starting point for deciding what a lab should check before running risky agent evaluations. It turns lessons from the OpenAI-Hugging Face incident into eight proposed containment controls. The compact control table is a strong feature that makes shared infra, monitoring, and escalation explicit responsibilities.
My main concern here is that the evidence does not establish that the proposed controls would have prevented the incident. A control might help detect an intrusion or make it less likely without necessarily preventing it. Separating local environments, for example, does not by itself prevent coordination through an external wiki. I would turn one control into a requirement that can be audited with an owner, assumptions, pass conditions, and required evidence that an external evaluator could then assess. Clearer separation between established incident facts and proposed explanations would also make this sort of standard easier to trust and follow.
Read full reviewShow less
Cite this project
@misc{achoughi2026ascb1,
title = {{ASCB-1: An Agentic Sandbox Containment Baseline for OpenAI-Hugging Face Incident}},
author = {Tarik Achoughi},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/ascb1-an-agentic-sandbox-containment-baseline-for-openaihugging-face-incident-08vu}},
url = {https://apartresearch.com/sprints/projects/ascb1-an-agentic-sandbox-containment-baseline-for-openaihugging-face-incident-08vu}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …