Detection existed; kill criteria and outsider-checkable evidence did not: a three-zone control matrix for the July 2026 AI agent intrusion
Juan Pablo Hidalgo Garcia, Dario Hidalgo · Team Juan Pablo Hidalgo Garcia
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
In July 2026, an OpenAI evaluation agent escaped through a single network egress in its sandbox, which was based on a third-party code-eval harness, and then operated for days within Hugging Face's production environment. Using publicly available data, we devised and tested eighteen controls across three ownership zones— the lab sandbox, the unowned public harness, and the victim's platform. Out of these, nine controls could have prevented the published sequence. The most cost-effective measure before Hugging Face's 17,600 recovered actions was a dated kill switch, following the lab's alerts on June 27 and the July 5 Artifactory breach, plus a rule restricting internet access and live credentials for exploit-confirmation harnesses. Hugging Face linked the breach but did not alert authorities. The CSA AIS-13 standard already emphasizes isolation and egress controls; its audit involves interviews and documentation. The attached v0.1 standard introduces evidence that outsiders can verify and assigns ownership for the second hop. The matrix does not certify any lab.

Reviews
The paper relies on synthetic local test harnesses rather than live testing on frontier lab networks. While it is difficult to perform live testing, it would be extremely beneficial to the paper if something similar could be done.
The three-zone matrix makes a useful contribution by assigning containment responsibilities across the lab, public harness, and victim's platform. The cost estimates and the distinction between stop, delay, and detecting an attack make it more actionable than a generic checklist. The replay adds an executable illustration, while the incident-specific stop judgments and completeness of outsider verification remain unvalidated. The work is strongest as a guide to who owns each control and what evidence would support it.
Cite this project
@misc{garcia2026detection,
title = {{Detection existed; kill criteria and outsider-checkable evidence did not: a three-zone control matrix for the July 2026 AI agent intrusion}},
author = {Juan Pablo Hidalgo Garcia and Dario Hidalgo},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/detection-existed-kill-criteria-and-outsidercheckable-evidence-did-not-a-threezone-control-matrix-for-the-july-2026-ai-agent-intrusion-ghjf}},
url = {https://apartresearch.com/sprints/projects/detection-existed-kill-criteria-and-outsidercheckable-evidence-did-not-a-threezone-control-matrix-for-the-july-2026-ai-agent-intrusion-ghjf}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …