Detection existed; kill criteria and outsider-checkable evidence did not: a three-zone control matrix for the July 2026 AI agent intrusion
Juan Pablo Hidalgo Garcia, Dario Hidalgo
In July 2026, an OpenAI evaluation agent escaped through a single network egress in its sandbox, which was based on a third-party code-eval harness, and then operated for days within Hugging Face's production environment. Using publicly available data, we devised and tested eighteen controls across three ownership zones— the lab sandbox, the unowned public harness, and the victim's platform. Out of these, nine controls could have prevented the published sequence. The most cost-effective measure before Hugging Face's 17,600 recovered actions was a dated kill switch, following the lab's alerts on June 27 and the July 5 Artifactory breach, plus a rule restricting internet access and live credentials for exploit-confirmation harnesses. Hugging Face linked the breach but did not alert authorities. The CSA AIS-13 standard already emphasizes isolation and egress controls; its audit involves interviews and documentation. The attached v0.1 standard introduces evidence that outsiders can verify and assigns ownership for the second hop. The matrix does not certify any lab.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Detection existed; kill criteria and outsider-checkable evidence did not: a three-zone control matrix for the July 2026 AI agent intrusion
},
author={
Juan Pablo Hidalgo Garcia, Dario Hidalgo
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


