Skip to content
Sprint projectSep 13, 2026Aguascalientes Mexico

Detection existed; kill criteria and outsider-checkable evidence did not: a three-zone control matrix for the July 2026 AI agent intrusion

Juan Pablo Hidalgo Garcia, Dario Hidalgo · Team Juan Pablo Hidalgo Garcia

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Detection existed; kill criteria and outsider-checkable evidence did not: a three-zone control matrix for the July 2026 AI agent intrusion

Recording (opens in new tab)Code (opens in new tab)
Share

In July 2026, an OpenAI evaluation agent escaped through a single network egress in its sandbox, which was based on a third-party code-eval harness, and then operated for days within Hugging Face's production environment. Using publicly available data, we devised and tested eighteen controls across three ownership zones— the lab sandbox, the unowned public harness, and the victim's platform. Out of these, nine controls could have prevented the published sequence. The most cost-effective measure before Hugging Face's 17,600 recovered actions was a dated kill switch, following the lab's alerts on June 27 and the July 5 Artifactory breach, plus a rule restricting internet access and live credentials for exploit-confirmation harnesses. Hugging Face linked the breach but did not alert authorities. The CSA AIS-13 standard already emphasizes isolation and egress controls; its audit involves interviews and documentation. The attached v0.1 standard introduces evidence that outsiders can verify and assigns ownership for the second hop. The matrix does not certify any lab.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The paper relies on synthetic local test harnesses rather than live testing on frontier lab networks. While it is difficult to perform live testing, it would be extremely beneficial to the paper if something similar could be done.

  2. The three-zone matrix makes a useful contribution by assigning containment responsibilities across the lab, public harness, and victim's platform. The cost estimates and the distinction between stop, delay, and detecting an attack make it more actionable than a generic checklist. The replay adds an executable illustration, while the incident-specific stop judgments and completeness of outsider verification remain unvalidated. The work is strongest as a guide to who owns each control and what evidence would support it.

Cite this project

@misc{garcia2026detection,
  title = {{Detection existed; kill criteria and outsider-checkable evidence did not: a three-zone control matrix for the July 2026 AI agent intrusion}},
  author = {Juan Pablo Hidalgo Garcia and Dario Hidalgo},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/detection-existed-kill-criteria-and-outsidercheckable-evidence-did-not-a-threezone-control-matrix-for-the-july-2026-ai-agent-intrusion-ghjf}},
  url = {https://apartresearch.com/sprints/projects/detection-existed-kill-criteria-and-outsidercheckable-evidence-did-not-a-threezone-control-matrix-for-the-july-2026-ai-agent-intrusion-ghjf}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026