Skip to content
Sprint projectSep 13, 2026Hyderabad

Falsifying Isolation Claims From Outside the Lab

Pranjal Gulati · Team dezz

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Falsifying Isolation Claims From Outside the Lab

Code (opens in new tab)
Share

An admission gate that lets an outside party falsify a lab's "the evaluation environment is isolated" claim without any access to the lab's network.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The admission-gate framing is right, containment claims should be falsifiable from outside, and verifier-controlled canary endpoints are a concrete mechanism. 144-cell enumeration is thorough, and they are clear falsification is not proof. Problem: this overlaps a lot with the Declared Isolation entry, the two should probably be one suite. Execution is the weak link, synthetic and protocol-level, agent study has acknowledged confounds. So the mechanism is demonstrated, not validated. Next: run the gate against a real sandbox product.

  2. The fact that the asymmetry in the middle is correct and well reasoned is that no external observer can prove that a network path is missing, but a breakdown in isolation can be made known in a log which the operator is not allowed to write to. This is the only one paper in which the decisive evidence is not provided by the party being examined, which is exactly what the track sets out to achieve, and by refusing to allow any conclusion to state that isolation has been achieved the appropriate discipline is upheld. The record entry derived from policy is the most important single outcome since it identifies the honest but mistaken operator rather than the dishonest one, and this is precisely what the reported incidents describe. The repository backs up the claim that is most important: one command lists all 144 configurations and the 22 probe-schedule cases from the standard library alone, with no need for a network connection, a model or credentials, and the CI system runs the regression suite on two platforms, so your central figure can indeed be recomputed by anyone. There are three things to fix. The most obvious is the scope: the Hugging Face escape went through a zero-day via a permitted egress, so a harmless probe would not have triggered that path no matter how many times it repeated; make it clear that the method deals with misconfiguration, not with abuse of a permitted exit. Second, the 24-of-24 versus 16-of-24 comparison is nearly true by design, so its persuasiveness is less than its precision indicates. Third, the commitments audit is present in your code but never appears in the report — the classification and agreement mechanisms are there, so publish the results of the audit rather than just stating that they exist, and note that the saved runs are stored outside the repository, which means that the agent study in particular cannot be verified by a reader.

    Read full reviewShow less

Cite this project

@misc{gulati2026falsifying,
  title = {{Falsifying Isolation Claims From Outside the Lab}},
  author = {Pranjal Gulati},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/falsifying-isolation-claims-from-outside-the-lab-5o9l}},
  url = {https://apartresearch.com/sprints/projects/falsifying-isolation-claims-from-outside-the-lab-5o9l}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026