Seen Early, Told Late: Who Detects, Links and Discloses AI-Agent Containment Incidents
Yutai Lin
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
When a developer's AI agent acts outside its sandbox, who notices, who names the developer, and how long does that take? We built a ledger of public cases up to 12 September 2026: 13 incident families and 31 targets, with 283 script-checked quotes. In seven fully dated families, attribution took 10 and 12 days in two and 74 to 129 days in five; the two fast cases began last. Victims noticed first at four targets but could not say whose agents they were. Thirteen of fourteen outsider attributions rested, directly or indirectly, on agents naming their own lab. Publication tracked the calendar: day-dated links made after Hugging Face's 16 July disclosure were published within eight days, while earlier developer-only links waited at least five weeks. We argue reporting clocks should start at a developer's earliest internal record of its agents' activity, and pose seven dated questions, four testing the account.
Reviews
the MIT AI Incident Tracker is also relevant work in this domain. I like this as research into previous incidents, but it is unclear to me how this ledger offers value above existing tools. It might be worth asking government bodies/regulators what they think is missing from existing repositories that could make their lives easier.
The project is a substantive synthesis of 31 targets across 13 incident families, tracing who detects an incident, links it to a developer, and discloses it. It documents how early detection can coexist with delayed attribution and publication, providing a useful empirical foundation for improving AI incident coordination and reporting.
Strength: Separating detection, attribution, and disclosure reveals a coordination problem that incident counts alone miss: a victim can notice harmful activity quickly without knowing which developer can help stop it. The cross-case chronology is a valuable foundation for AI incident governance.
Recommendation:
- Strengthen the explanations built from the chronology. Distinguish observed patterns from the mechanisms proposed to explain them. For example, labels appearing in 13/14 attribution chains do not establish that attribution would fail without those labels; network evidence may also identify the developer. Comparing the different routes to attribution would clarify which bottlenecks an intervention should address.
- Connect the reporting proposal to a plausible reduction in harm. Work through how an early internal alert would lead to a preliminary notice and a concrete response by an affected party. The May 26 observation and June 27 alert provide useful examples. This would make the policy case rest on improved response capability as well as earlier disclosure.
Read full reviewShow less
This project makes a useful contribution by separating when suspicious agent activity was noticed, when it was linked to a developer, and when it became public. It builds a timeline across public incidents and makes disputed dates and corrections visible, helping readers judge what the available evidence actually establishes.
The stronger conclusions depend on when organizations understood what had happened, which the dates do not always establish. A visit from a developer-associated address, for example, does not show that someone recognized an incident. With so few cases, I would avoid drawing firm conclusions about incentives to disclose. The timeline is useful, but it needs to distinguish recorded activity from evidence that people knew about the incident.
Cite this project
@misc{lin2026seen,
title = {{Seen Early, Told Late: Who Detects, Links and Discloses AI-Agent Containment Incidents}},
author = {Yutai Lin},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/seen-early-told-late-who-detects-links-and-discloses-aiagent-containment-incidents-eyj0}},
url = {https://apartresearch.com/sprints/projects/seen-early-told-late-who-detects-links-and-discloses-aiagent-containment-incidents-eyj0}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …