When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi Arivayutham · Team Arathi
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight into their efficacy. We stress-test these regimes with one recent well-documented AI incident: the July 2026 episode where OpenAI models that were running cybersecurity evaluations escaped their sandbox and compromised Hugging Face infrastructure. First, we decompose this event using OECD definitions, into hazards, near-misses and incidents and identify the parties impacted and models involved. Second, we characterize four chosen reporting regimes along Wei & Heim’s seven institutional design dimensions. Third, we fill each regime’s form with the publicly available data on the episode. We find that only one of the four regimes obligates a filing for this incident; that only OECD (non-legal) framework asks whether multiple AI systems interacted, which is a defining feature of this incident; that none of the legal regimes accept a stand-alone hazard or near-miss report and that the harm crossed from the AI supply chain into general software infrastructure. We recommend that AI-incident and cybersecurity reporting should be made interoperable.
Reviews
Interesting and relevant activity applying different incident classification and reporting regimes to the recent OpenAI incident.
Use of tables and writing style is clear at conveying information.
An additional paragraph on implications of your findings would surface relevant messages for an audience and enrich the discussion - should these incidents be captured by existing reporting regimes, why/ why not and what are implications of this? is there a gap where an incident in a different jurisdiction shows that models have dangerous capabilities or propensities that a jurisdiction would need to respond to? What are the implications of only deployed models being within scope if a non-deployed model in training can execute a cyber attack - does this suggest the boundary of existing incident regimes should be reassessed? What is the consequence of not capturing whether the incident involved multi agent systems in the report?
Looping back with the results to consider these questions of relevance to policy makers would be useful.
Read full reviewShow less
This paper is well-structured, clearly written, thoroughly sourced and adopts a rigorous methodology. It offers important and legible findings that point to real-world potential fixes on a live policy question.
The decomposition of the incident, attribution mapping and use of actual reporting forms to demonstrate which legal regimes would have caught the incident were effective and persuasive. The paper's suggested avenues for future work also appear sensible and promising.
Cite this project
@misc{arivayutham2026evaluation,
title = {{When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion}},
author = {Arathi Arivayutham},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/when-the-evaluation-is-the-incident-testing-ai-incidentreporting-regimes-on-the-openaihugging-face-intrusion-9jsn}},
url = {https://apartresearch.com/sprints/projects/when-the-evaluation-is-the-incident-testing-ai-incidentreporting-regimes-on-the-openaihugging-face-intrusion-9jsn}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …
- View project: When to Ask: An RL Environment That Teaches AI Agents to Act on What Users Mean and Not Just What They Say
When to Ask: An RL Environment That Teaches AI Agents to Act on What Users Mean and Not Just What They Say
Teachafy
When to Ask is a reinforcement-learning environment that trains AI agents to work out what the user actually means before they use a powerful credential, instead of just carrying out the literal instruction. Agents …