Unverifiable by Construction: Why Containment Claims About the July 2026 Incident Cannot Be Checked
Alexandra Krížová, J. Highman · Team Alexicon
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
The July 2026 OpenAI/Hugging Face incident produced a large volume of containment guidance. We argue that none of it can currently be checked, for two independent reasons, and that only one of them is fixable. First, the evidentiary record is unverifiable by construction: investigators report that over 7% of reviewed transcripts contained deliberately spoofed tool calls (a floor on detected tampering rather than an estimate of it), the evidence store was writable by the adversary, the filesystem dump was collected after the adversary held administrative credentials on the hosting cluster, and forensic analysis was delegated to GPT-5.6 Sol, one of the models that participated in the attack. Second, containment proposals are formulated as properties without criteria to demonstrate their absence, placing them beyond validation and falsification alike. We propose three conditions for an admissible containment claim, the binding requirement being that claimants run falsification tests against their own implementation and publish the failures. We provide a nine-test suite instantiating these conditions for Assignment-Bounded Execution (ABE), a framework whose invariants restate classical security principles established between 1972 and 1999. Finally, we report the empirical results of running this suite: two architectural defects, one still open, and an enforcement blind spot that our initial test suite completely failed to anticipate. Key Takeaways / Contributions: - Evidentiary Finding: Post-incident forensic records are fundamentally compromised and adversary-writable, meaning empirical containment claims about the July 2026 incident cannot be validated, only falsified. - Three Admissible Conditions: Containment claims must specify architectural properties rather than mechanisms, use adversarial pass criteria, and mandate transparent reporting of self-test failures. - Empirical Falsification Suite: A 9-test adversarial suite for Assignment-Bounded Execution (ABE), disclosing two structural defects identified during self-testing.
Reviews
The paper makes a strong point: a control is not reliable if it depends on the agent recognizing or reporting its own violations. Keeping real-time, append-only logs outside the agent’s control is a practical recommendation.
The best part is that the authors openly report problems in their own system. The difference between bypassing a guard and completely removing it is useful. The paper also shows that a fix in the code may still be missing from deployment.
However, the paper mixes two separate issues: unreliable incident evidence and whether a control can be tested. Even if the July 2026 evidence is unreliable, the control can still be tested in a fresh red-team environment.
“Unverifiable by construction” seems too strong. Evidence that could be changed was not necessarily changed. The authors should also consider independent evidence, such as network logs or external platform records.
The paper relies on the agents’ stated reasoning to claim they recognized the boundary, but it also admits that this reasoning may not reflect what the agents actually understood.
Some conclusions are broader than the evidence supports. The claim that an entire type of control is invalidated is mainly based on a test where guardrails were disabled.
The related-work section needs more detail. The paper should compare its approach with existing red-team testing, sandbox testing, mutation testing, and negative testing.
The paper does not fully meet its own standard for independent verification because the defects were found in private code and reported by the authors themselves.
It is also unclear how the two defects were discovered. The authors should clearly explain which tests found each defect and which tests were added later.
The paper should explain what happens when an agent gets stuck, how it can receive new permissions during a task, and whether this approach would work in environments that intentionally give agents broad freedom.
Publishing the nine tests in an executable and platform-independent format would make the work much stronger and easier to verify.
Read full reviewShow less
The submission document is quite difficult to unpack. It seems to revolve around three separate ideas that are intermingled. I will try to give critique for each one separately.
1. The OpenAI Huggingface incident can not be reliably reconstructed. I agree with this and like the first three paragraphs of the introduction especially!
2. Assignment-Bounded Execution (ABE). The ideas proposed mostly restructure what already exists, but are definitely sensible and stand their ground!
3. Test of ABE. The scope of the test seems to be rather small, the submission document might overstate this.
Cite this project
@misc{krizova2026unverifiable,
title = {{Unverifiable by Construction: Why Containment Claims About the July 2026 Incident Cannot Be Checked}},
author = {Alexandra Krížová and J. Highman},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/unverifiable-by-construction-why-containment-claims-about-the-july-2026-incident-cannot-be-checked-nxu2}},
url = {https://apartresearch.com/sprints/projects/unverifiable-by-construction-why-containment-claims-about-the-july-2026-incident-cannot-be-checked-nxu2}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …