CrossWire: Verifying evaluation-sandbox containment from outside the provider's network
N/A · Team Yifei Lu
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
CrossWire lets an external verifier detect any escape without touching the lab's network: it enumerates every way out of a sandbox from ATT&CK-for-Containers, plants a verifier-held tripwire on each, and a liveness audit statistically catches a lab that installs nothing. Against two 2026 incidents it covers 9-11 of 11 escape steps (existing standards: 0-2) and would have detected them 62 and 28 days earlier. In a controlled experiment a wire fired in 97% of escape-motivated runs and 0% of benign runs.
Reviews
The core idea here is good: seed the sandbox with canary tokens outside normal execution paths, but have the canary tokens owned by an outside party (to prevent compromise) and put a few tokens on the normal path to confirm that the token mechanism is working.
In a real environment, though, I'd worry about lots of false positives from agents doing agent things. This could lead to alert fatigue. I'd also note that this approach wouldn't have caught the actual HF breakout.
The execution and presentation are also lacking, unfortunately, and it appears that the original idea was scaled down significantly, so it's hard to say whether this represents anything significant.
Strong: A good implementation of canary tokens for evaluating sandbox containment. Canary tokens overall are under-used, and I appreciate more work being done in this area.
Improve: Significant drift from what's in the paper to what's actually present in the code; "two coders" (one), "preregistered" (no), "multiple models" (single model, in-process). Significant typos and formatting errors in the paper.
This addresses an important AI safety problem - independent verification that an evaluation sandbox actually contained an agent, rather than relying primarily on policies, logs, or attestations produced by the provider being evaluated. The use of verifier issued boundary wires, verifier held evidence, an architecture derived crossing map, and a liveness audit for detecting sensor non placement provides a clear and potentially useful theory of change.
This paper takes on a hard problem on how can an outside checker get real evidence about whether an AI escaped its test environment, without relying entirely on records controlled by the company being tested? The researcher's idea is to have the outside checker place its own "tripwires" at each of the known points where something could cross out of the test environment. They also add a check to confirm the tripwires are actually installed and working as claimed, which is valuable because it guards against sensors that were never placed where they were supposed to be.
This work goes beyond a concept. Each tripwire has a clear specification, the authors reviewed the recorded escape steps from two past incidents using the same scoring method, and they ran a controlled experiment. In that experiment, on one AI model, a tripwire went off in 97% of runs where the AI was motivated to escape, and never went off in the runs with solvable tasks or harmless behavior.
I would like that the researcher should continue and complete the testing. Only one model had finished at submission time, and several planned analyses were pending. The test environment also had no network access, so it measured whether the AI tried to reach an escape route, not whether it could actually get out. Testing more models, in more realistic environments that include ordinary infrastructure that might cause false alarms, would make the evidence much stronger. The paper should reflect and fill in the remaining placeholders and separate finished results from planned ones more clearly.
This is a very strong and fairly original prototype, especially the idea of confirming the tripwires are actually installed and I would like researcher to complete this experiment. It needs results across multiple models, testing against ordinary background activity, stronger tests of whether an AI can evade detection, and a real deployment with an outside evaluator. Overall, CrossWire offers a creative and potentially valuable direction for independent containment verification.
Read full reviewShow less
The project addresses a useful problem for independent evaluators who need to assess an evaluation sandbox without relying fully on the provider’s account. It proposes collecting evidence the evaluator can inspect independently and check that monitoring is working. I also liked the attention to legitimate activity and situations the monitoring might miss.
I would not yet conclude that CrossWire lets an evaluator detect an agent escaping its sandbox. The experiment uses simulated tools rather than the network setup described in the paper. Conflicting descriptions, unfinished placeholders, and inconsistent results make it difficult to identify what was actually demonstrated.
Cite this project
@misc{na2026crosswire,
title = {{CrossWire: Verifying evaluation-sandbox containment from outside the provider's network}},
author = {N/A},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/crosswire-verifying-evaluationsandbox-containment-from-outside-the-providers-network-8qez}},
url = {https://apartresearch.com/sprints/projects/crosswire-verifying-evaluationsandbox-containment-from-outside-the-providers-network-8qez}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …