When Our System Crosses the Line: An Organizational Preparedness Self-Check for AI Boundary-Crossing Incidents
Xiaochuan Wang · Team 706
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
This is Xiaochuan's independent entry (with Yiran Huang of Zhejiang Gongshang University) in the Apart Research AI Incident Response Sprint, a research hackathon running 11–13 September 2026, built on the Hugging Face July 2026 security incident. The current deliverable is Our System Went Out of Bounds — an organizational preparedness self-check: a paper-based tabletop exercise kit for small, resource-poor organizations. Its core claim is that the method does not produce verdicts but helps an organization build a risk structure for after failure: six control points and a three-layer model (reachability / cascade / stake points), run on a still-in-development human-machine World Engine front end, supported by 25 citations across four disciplines.
Reviews
The audience is real and underserved — organizations adopting AI through a vendor, where permissions accumulate during deployment and nobody keeps track of what the system can reach. Treating "don't know" as the valuable answer rather than a failure is the right instinct for that setting.
The difficulty is that the paper describes a package the submission does not contain. The figure generation scripts described as shipping do not ship, so none of the figure's numbers can be checked by a reader. The workshop pack said to be what this submission ships — handbook, facilitator's manual, forms, appendices — is absent, and the body repeatedly defers its actual operational content to those missing appendices. The assessment engine at the centre of the design is written about throughout in the present tense, as something that holds state and computes outcomes, while the only status given for it is that it remains in development.
Most seriously, one section reports that most organizations completing the grid for the first time return more than half their answers as "don't know" — while the results section states the package has no measured data and no subjects, and the abstract says it has never been piloted. That reads as a finding from sessions that never took place, and a reader who notices will discount work that does not deserve it. Either source that claim or cut it; it is the single most important change to make before this goes any further.
One session with one real organization, reported honestly, would be worth more than the theoretical apparatus currently carrying the argument. On the writing: the borrowed frameworks supporting a short checklist are more scaffolding than the idea needs. Cutting to the two or three that genuinely do work, and rewriting in the plain register the intended readers actually use, would make the package far more likely to be picked up by the organizations it is built for.
Read full reviewShow less
Cite this project
@misc{wang2026our,
title = {{When Our System Crosses the Line: An Organizational Preparedness Self-Check for AI Boundary-Crossing Incidents}},
author = {Xiaochuan Wang},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/when-our-system-crosses-the-line-an-organizational-preparedness-selfcheck-for-ai-boundarycrossing-incidents-dv52}},
url = {https://apartresearch.com/sprints/projects/when-our-system-crosses-the-line-an-organizational-preparedness-selfcheck-for-ai-boundarycrossing-incidents-dv52}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …