Skip to content
Sprint projectSep 14, 2026Toronto

When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion

Arathi Arivayutham · Team Arathi

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion

Share

AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight into their efficacy. We stress-test these regimes with one recent well-documented AI incident: the July 2026 episode where OpenAI models that were running cybersecurity evaluations escaped their sandbox and compromised Hugging Face infrastructure. First, we decompose this event using OECD definitions, into hazards, near-misses and incidents and identify the parties impacted and models involved. Second, we characterize four chosen reporting regimes along Wei & Heim’s seven institutional design dimensions. Third, we fill each regime’s form with the publicly available data on the episode. We find that only one of the four regimes obligates a filing for this incident; that only OECD (non-legal) framework asks whether multiple AI systems interacted, which is a defining feature of this incident; that none of the legal regimes accept a stand-alone hazard or near-miss report and that the harm crossed from the AI supply chain into general software infrastructure. We recommend that AI-incident and cybersecurity reporting should be made interoperable.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Interesting and relevant activity applying different incident classification and reporting regimes to the recent OpenAI incident.

    Use of tables and writing style is clear at conveying information.

    An additional paragraph on implications of your findings would surface relevant messages for an audience and enrich the discussion - should these incidents be captured by existing reporting regimes, why/ why not and what are implications of this? is there a gap where an incident in a different jurisdiction shows that models have dangerous capabilities or propensities that a jurisdiction would need to respond to? What are the implications of only deployed models being within scope if a non-deployed model in training can execute a cyber attack - does this suggest the boundary of existing incident regimes should be reassessed? What is the consequence of not capturing whether the incident involved multi agent systems in the report?

    Looping back with the results to consider these questions of relevance to policy makers would be useful.

    Read full reviewShow less
  2. This paper is well-structured, clearly written, thoroughly sourced and adopts a rigorous methodology. It offers important and legible findings that point to real-world potential fixes on a live policy question.

    The decomposition of the incident, attribution mapping and use of actual reporting forms to demonstrate which legal regimes would have caught the incident were effective and persuasive. The paper's suggested avenues for future work also appear sensible and promising.

Cite this project

@misc{arivayutham2026evaluation,
  title = {{When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion}},
  author = {Arathi Arivayutham},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/when-the-evaluation-is-the-incident-testing-ai-incidentreporting-regimes-on-the-openaihugging-face-intrusion-9jsn}},
  url = {https://apartresearch.com/sprints/projects/when-the-evaluation-is-the-incident-testing-ai-incidentreporting-regimes-on-the-openaihugging-face-intrusion-9jsn}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026