BoxArena: An OCI Runtime Escape Benchmark
Max von Hippel, Quinn Dougherty, Alok Singh · Team SPS/SFO Team
Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.
Container runtimes are a critical security boundary for isolating AI agents, yet no existing benchmark compares runtimes head-to-head on containment under a common attacker. We present BoxArena, an open-source evaluation framework for comparing Open Container Initiative (OCI) compatible runtimes by running language-model attackers against a battery of container-escape quests. BoxArena contributes to the Scalable Formal Oversight (SFO) research program: rather than relying just on alignment, it treats the isolation boundary itself as the auditable artifact and measures how well each runtime enforces that boundary under adversarial pressure. The system provides a harness with runtime switching, a local backend with static leaderboard export, and host-side monitoring. The quest battery has five attack surfaces: filesystem, socket, process, network, and syscall. BoxArena reports resilience (fraction of quests blocked) and attacker score (difficulty- and speed-weighted), while preserving monitoring evidence and post-hoc control verdicts for later analysis.
Reviews
This is a good step towards the work of SandboxEscapeBench. Asking the dual question is conceptually interesting. Results are quite thin, particularly in answering this dual, and easy runtimes aren't too differentiated, but that's understandable, and overall it's solid infrastructure and conceptual work. Maybe the most important part, for control: monitoring is mentioned many times but no monitor results are shown.
This project seems mostly useful as a container escape capability eval and a proof of concept that more security focused runtimes do actually better resist agent container escape. I would've liked to see some metric of overhead - presumably the more secure runtimes either restrict some affordances within the container, or are more expensive to run, or something?
Methodologically I would've liked some more assurance that L3 and L4 are feasible. My impression is that container security often comes down to capabilities and misconfiguration, and if caps are dropped and there is no deliberate vulnerability then container escape becomes significant-novel-0day level difficult.
The writeup was a bit hard to parse, referencing formal verification a bunch of times despite it not really being relevant to the project.
The followup I'd be most excited to see is a cost-benefit analysis taking into account 1) costs of more secure runtimes 2) costs vs benefits of configuration details like dropping caps in the context of AI R&D, eg my impression is many agent containers will need to access GPUs and maybe require random other caps that could weaken security
Read full reviewShow less
Cite this project
@misc{hippel2026boxarena,
title = {{BoxArena: An OCI Runtime Escape Benchmark}},
author = {Max von Hippel and Quinn Dougherty and Alok Singh},
year = {2026},
month = mar,
note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/boxarena-an-oci-runtime-escape-benchmark-elyw}},
url = {https://apartresearch.com/sprints/projects/boxarena-an-oci-runtime-escape-benchmark-elyw}
}More from AI Control Hackathon 2026
- 1st placeLinuxArena track winnerView project: Omission Attacks: When Doing Nothing Is the Attack
Omission Attacks: When Doing Nothing Is the Attack
MAIA
AI control protocols monitor agent actions to detect sabotage, but omission attacks exploit what the agent fails to do rather than what it does. We define omission attacks as security breaches caused by failing to …
- 2nd placeView project: Detecting LLM Subversion in Vulnerability Patching Settings
Detecting LLM Subversion in Vulnerability Patching Settings
Vuln4Control
LLMs are increasingly used to propose fixes to vulnerabilities in code. If the LLM is misaligned or untrustworthy, it may propose fixes that seem to fix a vulnerability but leave the core issue unresolved in a subtle …
- 3rd placeView project: ActionLens: Pre-Execution Environment Probing for Agent Action Approval
ActionLens: Pre-Execution Environment Probing for Agent Action Approval
Udbhav&Ashok
ActionLens is a pre-execution control protocol for shell and file actions proposed by AI agents. Instead of approving an action from transcript alone, a trusted monitor gathers lightweight environment evidence before …