Falsifying Isolation Claims From Outside the Lab
Pranjal Gulati · Team dezz
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
An admission gate that lets an outside party falsify a lab's "the evaluation environment is isolated" claim without any access to the lab's network.
Reviews
The admission-gate framing is right, containment claims should be falsifiable from outside, and verifier-controlled canary endpoints are a concrete mechanism. 144-cell enumeration is thorough, and they are clear falsification is not proof. Problem: this overlaps a lot with the Declared Isolation entry, the two should probably be one suite. Execution is the weak link, synthetic and protocol-level, agent study has acknowledged confounds. So the mechanism is demonstrated, not validated. Next: run the gate against a real sandbox product.
The fact that the asymmetry in the middle is correct and well reasoned is that no external observer can prove that a network path is missing, but a breakdown in isolation can be made known in a log which the operator is not allowed to write to. This is the only one paper in which the decisive evidence is not provided by the party being examined, which is exactly what the track sets out to achieve, and by refusing to allow any conclusion to state that isolation has been achieved the appropriate discipline is upheld. The record entry derived from policy is the most important single outcome since it identifies the honest but mistaken operator rather than the dishonest one, and this is precisely what the reported incidents describe. The repository backs up the claim that is most important: one command lists all 144 configurations and the 22 probe-schedule cases from the standard library alone, with no need for a network connection, a model or credentials, and the CI system runs the regression suite on two platforms, so your central figure can indeed be recomputed by anyone. There are three things to fix. The most obvious is the scope: the Hugging Face escape went through a zero-day via a permitted egress, so a harmless probe would not have triggered that path no matter how many times it repeated; make it clear that the method deals with misconfiguration, not with abuse of a permitted exit. Second, the 24-of-24 versus 16-of-24 comparison is nearly true by design, so its persuasiveness is less than its precision indicates. Third, the commitments audit is present in your code but never appears in the report — the classification and agreement mechanisms are there, so publish the results of the audit rather than just stating that they exist, and note that the saved runs are stored outside the repository, which means that the agent study in particular cannot be verified by a reader.
Read full reviewShow less
Cite this project
@misc{gulati2026falsifying,
title = {{Falsifying Isolation Claims From Outside the Lab}},
author = {Pranjal Gulati},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/falsifying-isolation-claims-from-outside-the-lab-5o9l}},
url = {https://apartresearch.com/sprints/projects/falsifying-isolation-claims-from-outside-the-lab-5o9l}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …