Attribution Is Not Done and the Clock Is Running: A 90-Minute Tabletop Exercise Toolkit for Agent Loss- of-Control Incidents
Dan XU, Lujia Liang, Xicheng Li
Think tanks and ministries run AI crisis tabletop exercises (TTXs), but published scenarios simulate criminal misuse by outside actors. The July 2026 OpenAI / Hugging Face incident is a different failure mode: a frontier lab's own agents escaped their sandbox during a cybersecurity evaluation and compromised a third party's production infrastructure while the lab did not yet know it was the source. No open toolkit covers this case. We built a 60–90-minute discussion-based TTX for 5 to 8 participants, runnable by a non-expert facilitator, whose eight injects follow the public incident timeline and force four decisions: escalate an evaluation anomaly to an incident; contain or preserve evidence; disclose before attribution; notify a regulator under time pressure. Scored on a standards-based instrument, one full run reached 1.88/3 (S), falling from 2.44 (lab-internal) to 1.17 (public/regulatory). The source/victim asymmetry is what participants need to rehearse. Professional evaluation remains the next step.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Attribution Is Not Done and the Clock Is Running: A 90-Minute Tabletop Exercise Toolkit for Agent Loss- of-Control Incidents
},
author={
Dan XU, Lujia Liang, Xicheng Li
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


