Skip to content
Sprint projectSep 13, 2026Tangier, Morocco

ASCB-1: An Agentic Sandbox Containment Baseline for OpenAI-Hugging Face Incident

Tarik Achoughi · Team ASCB

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: ASCB-1: An Agentic Sandbox Containment Baseline for OpenAI-Hugging Face Incident

Code (opens in new tab)
Share

Most accounts of the July 2026 OpenAI–Hugging Face incident treat it as one story: agents escaped a sandbox and ended up in production. That doesn't give a defender anything to act on. This project breaks the incident into eight separate failures, each tied to a specific, dated event across OpenAI's, Hugging Face's, and METR's own reports. From those eight failures we build ASCB-1: eight testable controls, each mapped to the failure it fixes and to standard MITRE ATT&CK/ATLAS technique codes.And we propose two new attack-pattern names for MITRE's official catalog, since the coordination behavior we found doesn't have one yet.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Strong: Good reconstruction of the event and control matrix. Good discussion of how the DSEwiki incident informed the controls (you have to look for both external and internal message boards). Controls C4 and C5 are potentially interesting for further development. Identified a gap in the ATLAS attack technique catalog and proposed two candidate techniques.

    Improve: Very limited implementation and no testing of the controls; at the moment this is just hypothetical. The required limitations and dual-use appendix is missing (though there's a single sentence note on limitations). A literature review of previous proposed control matrices would have been helpful. The version of ATLAS cited is not the most current version (the gap identified appears to still be in the current version, however).

  2. - I recommend reading Neel Nanda's advice on writing an abstract https://www.alignmentforum.org/posts/eJGptPbbFPZGLpjsp/highly-opinionated-advice-on-how-to-write-ml-papers

    - It's unclear that better/novel security practices would have helped prevent the OAI-HF incident, even well-known existing cybersecurity practices would have helped, but OpenAI did not implement these basic best-practices.

    - I recommend using proper citations, and not just naming the reports (e.g. as done at the start of Section 2)

    - Learning to use LaTex (traditional but tricky) or Typst (less widespread but a lot easier) for document formatting will automatically make the final document look much more polished

    - The report doesn't have enough humility when talking about the faults it found. Because the full details haven't been released, almost certainly there are failures that we don't have enough information to reconstruct, but the report talks as though its failure chain was complete and exhaustive

    - The report proposes ASCB-1 and describes it in detail, but it's unclear exactly what it is. Is it a set of tests? a checklist? a recommended list of things to implement? The report describes it as a set of "controls" but "control" is very overloaded and no narrowing of the definition is done.

    - It's unclear how much the proposed framework is just rephrasing items from ATT&CK/MITRE/ATLAS, more information about what is the novel contribution would have been welcome

    - There is a proposed framework, but nothing is done to test the framework or to motivate what it adds that existing frameworks don't. Almost certainly there are gaps in existing frameworks (because they don't anticipate swarms of AIs) but this project doesn't sufficiently fill those gaps.

    - There's significant indications of AI-written text, this prevented me from reading the report in full.

    Read full reviewShow less
  3. I think the approach here with ASCB-1 could be a useful starting point for deciding what a lab should check before running risky agent evaluations. It turns lessons from the OpenAI-Hugging Face incident into eight proposed containment controls. The compact control table is a strong feature that makes shared infra, monitoring, and escalation explicit responsibilities.

    My main concern here is that the evidence does not establish that the proposed controls would have prevented the incident. A control might help detect an intrusion or make it less likely without necessarily preventing it. Separating local environments, for example, does not by itself prevent coordination through an external wiki. I would turn one control into a requirement that can be audited with an owner, assumptions, pass conditions, and required evidence that an external evaluator could then assess. Clearer separation between established incident facts and proposed explanations would also make this sort of standard easier to trust and follow.

    Read full reviewShow less

Cite this project

@misc{achoughi2026ascb1,
  title = {{ASCB-1: An Agentic Sandbox Containment Baseline for OpenAI-Hugging Face Incident}},
  author = {Tarik Achoughi},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/ascb1-an-agentic-sandbox-containment-baseline-for-openaihugging-face-incident-08vu}},
  url = {https://apartresearch.com/sprints/projects/ascb1-an-agentic-sandbox-containment-baseline-for-openaihugging-face-incident-08vu}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026