Sentinel: The Declared Scope Is the Unit of Containment Evidence in Agentic Evaluations
Crish Nagarkar
In the 2026 containment failures at OpenAI, Anthropic and AISI, each evaluation's network scope was declared in prose or assumed, never checked against what the agent did, and detection came late from a side signal. Sentinel treats that scope, written as a machine-readable allowlist, as the unit of containment evidence a third party can verify: a mechanical Inspect scanner classifies each tool-call destination against the list and pages on the first out-of-scope host. No model is in the gate. On reconstructions of three public incidents it pages at the first non-allowlisted destination. On the 14,591-revision DSEWiki export, a record we did not write, a wildcard scope (any GET) is silent and a data-source-only scope pages at the first wiki host, then floods. On a live eight-sample Inspect run with an empty declared scope it is correctly silent on 42 tool calls. The headline is an ablation: a too-tight scope pages on permitted traffic; a wildcard is silent on the escape. Verifiability is a property of the pair (trace, allowlist). Appendix C is a filled disclosure record for the ExploitGym class. The scanner audits containment; it does not provide it.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Sentinel: The Declared Scope Is the Unit of Containment Evidence in Agentic Evaluations
},
author={
Crish Nagarkar
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


