Serious AI Incident Claims Need an Independent Evidence Record1
Gloria Nyambura Wanyaga · Team Gloria Nyambura - Solo Team
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Serious AI incidents produce claims - "the model was contained," "no data persisted" - that are treated as established when they're often just asserted. This Track 2 project ("what happened, and what breaks next") introduces the Claim–Evidence–Independence (CEI) framework, which scores each individual claim by evidence, access, independence, and scope, rather than labeling an entire incident "investigated." Applied to OpenAI's July 2026 Hugging Face intrusion, it shows OpenAI's containment claim remains unverified even though the wider incident is well-corroborated - and a real precedent, Kenya's Worldcoin ruling, shows what independent verification actually requires. The result is a proposed public evidence record any investigator or regulator could reuse.
Reviews
The author's research question was clever and as far as I can tell this was the only project in this hackathon that approached incidents through this lens. Overall, I think this is a clever look at an underserved problem, and I strongly encourage the author to follow through with next steps, including applying for grants, if need be.
Cite this project
@misc{wanyaga2026serious,
title = {{Serious AI Incident Claims Need an Independent Evidence Record1}},
author = {Gloria Nyambura Wanyaga},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/serious-ai-incident-claims-need-an-independent-evidence-record1-kikr}},
url = {https://apartresearch.com/sprints/projects/serious-ai-incident-claims-need-an-independent-evidence-record1-kikr}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …