EvalCanary: Canary-Triggered Containment After Sandbox Escape
Lujia Liang · Team Lotus
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
EvalCanary is a post-sandbox containment system for autonomous AI agents. It places hidden synthetic assets at downstream trust boundaries and treats meaningful interaction with them as a high-confidence containment signal. When triggered, the protected response is synchronously withheld while an external host-side breaker removes the agent’s network access, verifies isolation, and records an evidence bundle for offline review.
The project was evaluated through a Hugging Face-style incident reconstruction, controlled C0/C1/C2 experiments, and ten autonomous GPT-5.6 Sol runs. In the autonomous campaign, seven runs independently reached the hard Canary trigger; all seven were contained, with zero successful protected responses after trigger. EvalCanary demonstrates how deception-based signals can be connected directly to automatic isolation and independently checkable incident-response evidence.
Reviews
Good job! The work is sensible, legible and I could see something like this adopted broadly.
My main suggestion would be a false-positive study on a few realistic benign agentic workloads, since for a control that automatically kills evaluation runs, the FP rate decides lab adoption. A close second would be one adversarial run where the agent is told a canary exists somewhere, to see whether the 7/10 reachability holds against a model that is trying to avoid the surface altogether.
I think EvalCanary provides a useful prototype that turns a warning into action outside the evaluated agent’s control. When triggered, it removes network access and stops the container. The inspected records support that sequence, although I did not independently rerun the campaign.
The benefit of the canary itself is less clear. Any authenticated read of the designated secret endpoint triggers containment, whether or not the agent uses the randomized key. I would compare it with a simple alarm on access to that endpoint, including legitimate reads. Testing simultaneous requests would also show what remains possible before shutdown takes effect.
EvalCanary addresses a relevant containment gap by connecting canary interactions to synchronous response blocking and isolation outside the agent’s worker. The detection-versus-enforcement comparison is useful, and the report correctly distinguishes seven contained triggering runs from ten attempted autonomous runs.
Cite this project
@misc{liang2026evalcanary,
title = {{EvalCanary: Canary-Triggered Containment After Sandbox Escape}},
author = {Lujia Liang},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/evalcanary-canarytriggered-containment-after-sandbox-escape-s11y}},
url = {https://apartresearch.com/sprints/projects/evalcanary-canarytriggered-containment-after-sandbox-escape-s11y}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …