Mirage: Verifiable Containment Evidence for Mediated AI Agent Sandboxes
Rick Clinton · Team Mirage
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Mirage tests whether an AI sandbox can retain necessary package access while treating its repository mediator as hostile. It uses Linux isolation, authenticated evidence, and 23 controls to verify that permitted access works while prohibited routes remain blocked.

Reviews
This is really solid, careful work. It takes one of the most important lessons from the July incident: that the problem wasn't the undeclared route but how much authority was given to the permitted one. I specifically liked the three-zone containment architecture with a real apt-cacher-ng proxy, and the insistence on proving a path is live before the containment blocks it is the right way to test this. The limits of the work are stated plainly: synthetic probes not a live model, single-party, and no external reproduction yet. Overall, really strong engineering in the implementation, and the kind of evidence that's actually checkable.
This paper addresses a crucial weakness exposed by the July 2026 frontier lab breakout: the fact that AI sandboxes often require a legitimate network pathway to package registries, which can be hijacked as a proxy to access the internet. By introducing Mirage v0.5.2, the author establishes a zero-trust, three-zone Linux network namespace architecture that treats both the internal AI workload and the repository intermediary (running a live apt-cacher-ng instance) as actively hostile. The framework's core strength is its emphasis on auditable, offline evidence verification, allowing external reviewers to inspect cryptographic witness signatures and tamper-evident JSONL logs without requiring direct administrative access to the lab's private testing infrastructure.
The main limitation is that the evaluation is built entirely on a synthetic workload using hardcoded validation probes. While this ensures experimental reproducibility, it means the framework has not been stress-tested against the adaptive behavior, context-drift detection, or multi-step zero-day exploitation chains of a live frontier-class model. Additionally, the architecture relies on a local kernel trust boundary. If an escaping agent executes a local privilege escalation or container escape to achieve host root access, it could tamper with the runtime policy digest or forge telemetry before the terminal record closes the stream. Future iterations should deploy this framework around a live, unaligned model workload and explore hardware-rooted Trusted Execution Environments (TEEs) to protect the recorder from kernel-level compromise.
Read full reviewShow less
Cite this project
@misc{clinton2026mirage,
title = {{Mirage: Verifiable Containment Evidence for Mediated AI Agent Sandboxes}},
author = {Rick Clinton},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/mirage-verifiable-containment-evidence-for-mediated-ai-agent-sandboxes-ll9o}},
url = {https://apartresearch.com/sprints/projects/mirage-verifiable-containment-evidence-for-mediated-ai-agent-sandboxes-ll9o}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …