CapsuleVM: An Object-Capability Containment Standard, Verifiable Black-Box
jiangplus
The July 2026 sandbox escape — two evaluation agents leaving through “the one network egress their sandbox
permitted” via a zero-day in a package-registry proxy, escalating privileges, and exfiltrating benchmark solutions from
Hugging Face — is usually read as a patching failure. We read it as a model-of-authority failure: the sandbox was an
access-control-list (ACL) design in which the agent held a socket, an escalatable identity, and no structural coupling
between “safety classifiers off” and “authority down”. We submit CapsuleVM for Track 1 (Containment): (1) an objectcapability (OCap) evaluation framework, written before the incident, used as a containment standard — 9 red lines, 8
weighted dimensions, 19 red-team cases; (2) a black-box verification suite that a third party runs against a harness’s wire
endpoint with no access to the lab’s network, model, or evaluation content; and (3) a zero-dependency reference
implementation; four rounds of external adversarial
review) whose full test gate passes on native x86_64 Linux. We map each documented incident fact to the rule it trips and
the regression test that pins the structural counterpart. Main takeaway: a standard that forbids ambient authority — no
socket in the guest, empty capability list at start, secret-read and outbound-sink never in one subject, no side effect
without a durable record — is implementable today, and verifiable without entering the lab
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) CapsuleVM: An Object-Capability Containment Standard, Verifiable Black-Box
},
author={
jiangplus
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


