Playing Through an AI Incident
Sofia Gallego · Team Nefinia
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Little Agent Lab is a small browser game where you build a team of AI agents to complete a task while keeping a private item from leaking. Three short missions escalate from basic capability separation to a final one where two guard characters offer to watch the gate — only one actually does, and you find out by running the plan, not by reading a description. It's an early proof of concept aimed at giving people, especially kids, a playable feel for ideas like oversight and verification. Play it at sofiagallego.com/little-agent-lab.
Reviews
Helping people understand why incidents occur - the concepts of monitoring, safeguards and permissions in practice and why they might fail - is important and could help make aspects of AI risks and incidents more accessible to new audiences.
The paper could have a stronger link to potential impact by building on the idea it is communicating that 'a safeguard that is present but does not work' and extending this to when this might happen, particularly with respect to advanced AI systems which can find workarounds. So it could cover different failure modes - wrong setup, not actually applied, a system can subvert it intentionally.
It would also have been interesting to see a rationale relating to the choice of aspects included for this initial prototype, as they relate to what the author considers most important to communicate to a non technical audience about AI incidents.
This is a delightful idea that I'd love to see fleshed out more. As society grapples with helping children grow up with a technology that could theoretically destroy the world, I appreciate any project that attempts to make it easier to explain!
Cite this project
@misc{gallego2026playing,
title = {{Playing Through an AI Incident}},
author = {Sofia Gallego},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/playing-through-an-ai-incident-d1bo}},
url = {https://apartresearch.com/sprints/projects/playing-through-an-ai-incident-d1bo}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …