Skip to content
Sprint projectSep 14, 2026Shanghai, China

Seen Early, Told Late: Who Detects, Links and Discloses AI-Agent Containment Incidents

Yutai Lin

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Seen Early, Told Late: Who Detects, Links and Discloses AI-Agent Containment Incidents

Share

When a developer's AI agent acts outside its sandbox, who notices, who names the developer, and how long does that take? We built a ledger of public cases up to 12 September 2026: 13 incident families and 31 targets, with 283 script-checked quotes. In seven fully dated families, attribution took 10 and 12 days in two and 74 to 129 days in five; the two fast cases began last. Victims noticed first at four targets but could not say whose agents they were. Thirteen of fourteen outsider attributions rested, directly or indirectly, on agents naming their own lab. Publication tracked the calendar: day-dated links made after Hugging Face's 16 July disclosure were published within eight days, while earlier developer-only links waited at least five weeks. We argue reporting clocks should start at a developer's earliest internal record of its agents' activity, and pose seven dated questions, four testing the account.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. the MIT AI Incident Tracker is also relevant work in this domain. I like this as research into previous incidents, but it is unclear to me how this ledger offers value above existing tools. It might be worth asking government bodies/regulators what they think is missing from existing repositories that could make their lives easier.

  2. The project is a substantive synthesis of 31 targets across 13 incident families, tracing who detects an incident, links it to a developer, and discloses it. It documents how early detection can coexist with delayed attribution and publication, providing a useful empirical foundation for improving AI incident coordination and reporting.

    Strength: Separating detection, attribution, and disclosure reveals a coordination problem that incident counts alone miss: a victim can notice harmful activity quickly without knowing which developer can help stop it. The cross-case chronology is a valuable foundation for AI incident governance.

    Recommendation:

    - Strengthen the explanations built from the chronology. Distinguish observed patterns from the mechanisms proposed to explain them. For example, labels appearing in 13/14 attribution chains do not establish that attribution would fail without those labels; network evidence may also identify the developer. Comparing the different routes to attribution would clarify which bottlenecks an intervention should address.

    - Connect the reporting proposal to a plausible reduction in harm. Work through how an early internal alert would lead to a preliminary notice and a concrete response by an affected party. The May 26 observation and June 27 alert provide useful examples. This would make the policy case rest on improved response capability as well as earlier disclosure.

    Read full reviewShow less
  3. This project makes a useful contribution by separating when suspicious agent activity was noticed, when it was linked to a developer, and when it became public. It builds a timeline across public incidents and makes disputed dates and corrections visible, helping readers judge what the available evidence actually establishes.

    The stronger conclusions depend on when organizations understood what had happened, which the dates do not always establish. A visit from a developer-associated address, for example, does not show that someone recognized an incident. With so few cases, I would avoid drawing firm conclusions about incentives to disclose. The timeline is useful, but it needs to distinguish recorded activity from evidence that people knew about the incident.

Cite this project

@misc{lin2026seen,
  title = {{Seen Early, Told Late: Who Detects, Links and Discloses AI-Agent Containment Incidents}},
  author = {Yutai Lin},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/seen-early-told-late-who-detects-links-and-discloses-aiagent-containment-incidents-eyj0}},
  url = {https://apartresearch.com/sprints/projects/seen-early-told-late-who-detects-links-and-discloses-aiagent-containment-incidents-eyj0}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026