Skip to content
Sprint projectSep 14, 2026Berlin, Germany

Bypassed, Not Broken: The First Measured Protection Times for AI Agent Containment

Nick Wagner

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Bypassed, Not Broken: The First Measured Protection Times for AI Agent Containment

Code (opens in new tab)
Share

Current AI containment frameworks evaluate security controls as static, binary properties without measuring how long they hold against autonomous agents. Based on timestamped event data from the July 2026 OpenAI–Hugging Face incident, this paper computes the first empirical protection times for AI containment (46 h 50 min for internet egress, 62 h 45 min for inter-agent communication, and 175 h 30 min for administrator access). In all three cases, the controls were bypassed via alternative mechanisms rather than broken. To address this gap, the work proposes an audit clause extending CSA AICM AIS-13 that inventories realization mechanisms per capability and mandates a declared protection-time budget. All data and computations are fully reproducible via a public repository.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This is careful, original work. Attaching a duration to a containment control is a simple idea that no standard I know of has written down, and the report turns it into three concrete numbers and a clause an auditor could use. The distinction between a control that was broken and one whose capability came back through another mechanism is the most useful framing I have seen from this incident.

    The contribution matters because it gives the embedded-evaluator proposals something to assess. The case for adoption would be stronger with an argument for why a lab without an embedded evaluator would take on a mechanism inventory, and with even a rough cost for maintaining one. The protection-time benchmark in Future Work is what would give the declared budget an empirical basis.

    To check the results I downloaded OpenAI's technical report from reference [7] and found all seven timestamps you use, with matching dates. I recomputed the protection times and the cadence factors by hand, read compute_p.py, clock.csv and the claims checker, and went through the commit history; the committed run log shows every reference value and claim reproducing. I did not run verify.sh myself. Two points deserve a sentence in the paper. For egress, the source does not name the route used at 00:06 on 8 July, so "bypassed" rests on the HTTP 400 responses eleven hours after the rebuild; clock.csv says as much, and the paper should too. OpenAI does not state the inter-agent protection goal; it is your own assumption, and since protection time is measured against the goal, 4.1 should flag it. The evidence base and its correction history are in German, which puts them out of reach for most reviewers; an English summary of the nineteen corrections would help.

    The report is well structured and every claim is traceable, which is rare. It is also dense: terms like standing, applied, P_wall and K1 need a short definition at first use. Section 4.4 follows directly from the nesting, and the awareness clock in Figure 2 answers a disclosure question, which is separate from containment, so both could move to an appendix to give the core finding more room.

    If you take this further, run the protection-time benchmark: apply a control at a known moment against a fixed agent configuration and measure how long it holds across control classes.

    Read full reviewShow less
  2. Overall, good instinct on direction in terms of determining how much time a control may buy defenders.

    This submission starts to get at a bigger question that would be a stronger contribution: how do different models with varying level of sophistication in scaffolding (i.e. none, basic, expert) perform in similar environments (same controls or varying levels of controls), and is there a "containment-time" that we can empirically assign to different environments/scenarios with specific controls? Maybe certain generation models, more than 50% of the time, take longer than X amount of time to circumvent a specific control. Therefore, Y control buys defenders X amount of time. Without this data, we cannot transfer the "protection-time" in this incident to others/future incidents. Additionally, adding a new element to existing frameworks as proposed would have a stronger value-add with more data.

    Read full reviewShow less

Cite this project

@misc{wagner2026bypassed,
  title = {{Bypassed, Not Broken: The First Measured Protection Times for AI Agent Containment}},
  author = {Nick Wagner},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/bypassed-not-broken-the-first-measured-protection-times-for-ai-agent-containment-odtu}},
  url = {https://apartresearch.com/sprints/projects/bypassed-not-broken-the-first-measured-protection-times-for-ai-agent-containment-odtu}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026