Skip to content
Sprint projectSep 13, 2026Slovakia/USA

Unverifiable by Construction: Why Containment Claims About the July 2026 Incident Cannot Be Checked

Alexandra Krížová, J. Highman · Team Alexicon

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Unverifiable by Construction: Why Containment Claims About the July 2026 Incident Cannot Be Checked

Share

The July 2026 OpenAI/Hugging Face incident produced a large volume of containment guidance. We argue that none of it can currently be checked, for two independent reasons, and that only one of them is fixable. First, the evidentiary record is unverifiable by construction: investigators report that over 7% of reviewed transcripts contained deliberately spoofed tool calls (a floor on detected tampering rather than an estimate of it), the evidence store was writable by the adversary, the filesystem dump was collected after the adversary held administrative credentials on the hosting cluster, and forensic analysis was delegated to GPT-5.6 Sol, one of the models that participated in the attack. Second, containment proposals are formulated as properties without criteria to demonstrate their absence, placing them beyond validation and falsification alike. We propose three conditions for an admissible containment claim, the binding requirement being that claimants run falsification tests against their own implementation and publish the failures. We provide a nine-test suite instantiating these conditions for Assignment-Bounded Execution (ABE), a framework whose invariants restate classical security principles established between 1972 and 1999. Finally, we report the empirical results of running this suite: two architectural defects, one still open, and an enforcement blind spot that our initial test suite completely failed to anticipate. Key Takeaways / Contributions: - Evidentiary Finding: Post-incident forensic records are fundamentally compromised and adversary-writable, meaning empirical containment claims about the July 2026 incident cannot be validated, only falsified. - Three Admissible Conditions: Containment claims must specify architectural properties rather than mechanisms, use adversarial pass criteria, and mandate transparent reporting of self-test failures. - Empirical Falsification Suite: A 9-test adversarial suite for Assignment-Bounded Execution (ABE), disclosing two structural defects identified during self-testing.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The paper makes a strong point: a control is not reliable if it depends on the agent recognizing or reporting its own violations. Keeping real-time, append-only logs outside the agent’s control is a practical recommendation.

    The best part is that the authors openly report problems in their own system. The difference between bypassing a guard and completely removing it is useful. The paper also shows that a fix in the code may still be missing from deployment.

    However, the paper mixes two separate issues: unreliable incident evidence and whether a control can be tested. Even if the July 2026 evidence is unreliable, the control can still be tested in a fresh red-team environment.

    “Unverifiable by construction” seems too strong. Evidence that could be changed was not necessarily changed. The authors should also consider independent evidence, such as network logs or external platform records.

    The paper relies on the agents’ stated reasoning to claim they recognized the boundary, but it also admits that this reasoning may not reflect what the agents actually understood.

    Some conclusions are broader than the evidence supports. The claim that an entire type of control is invalidated is mainly based on a test where guardrails were disabled.

    The related-work section needs more detail. The paper should compare its approach with existing red-team testing, sandbox testing, mutation testing, and negative testing.

    The paper does not fully meet its own standard for independent verification because the defects were found in private code and reported by the authors themselves.

    It is also unclear how the two defects were discovered. The authors should clearly explain which tests found each defect and which tests were added later.

    The paper should explain what happens when an agent gets stuck, how it can receive new permissions during a task, and whether this approach would work in environments that intentionally give agents broad freedom.

    Publishing the nine tests in an executable and platform-independent format would make the work much stronger and easier to verify.

    Read full reviewShow less
  2. The submission document is quite difficult to unpack. It seems to revolve around three separate ideas that are intermingled. I will try to give critique for each one separately.

    1. The OpenAI Huggingface incident can not be reliably reconstructed. I agree with this and like the first three paragraphs of the introduction especially!

    2. Assignment-Bounded Execution (ABE). The ideas proposed mostly restructure what already exists, but are definitely sensible and stand their ground!

    3. Test of ABE. The scope of the test seems to be rather small, the submission document might overstate this.

Cite this project

@misc{krizova2026unverifiable,
  title = {{Unverifiable by Construction: Why Containment Claims About the July 2026 Incident Cannot Be Checked}},
  author = {Alexandra Krížová and J. Highman},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/unverifiable-by-construction-why-containment-claims-about-the-july-2026-incident-cannot-be-checked-nxu2}},
  url = {https://apartresearch.com/sprints/projects/unverifiable-by-construction-why-containment-claims-about-the-july-2026-incident-cannot-be-checked-nxu2}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026