Two Witnesses: An evidentiary coalition audit of AI-agent incident disclosure
Michelle Wanjiku Thuo · Team African Civic Trust (ACT)
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
AI-agent incidents scatter evidence across organizations an agent touches. I built a method for testing, fact by fact, whether one organization’s evidence is enough to establish a safety relevant claim or whether more than one is needed. Applied to 32 facts from two 2026 incidents involving the same developer, a Hugging Face intrusion and an undisclosed wiki misuse, 26 were single stakeholder sufficient. Two facts needed evidence from both organizations when they became public. One, attribution of the Hugging Face intrusion to its developer, is the clearest case. Neither Hugging Face’s disclosure nor the developer’s internal signal alone identified who was responsible then, although the developer’s later account is sufficient today. The other still needs both sides. This demonstrates the method, not that incidents generally need more than one witness or that environment or collective intelligence explains it. The companion tool, Evidence Coalition Explorer, lets readers test all 32 facts.
Reviews
A neat idea: does one company's evidence prove a fact alone, or do you need both sides together? Applied to 32 real facts across two 2026 incidents - 26 needed just one side, 2 needed both, with attribution of the Hugging Face intrusion as the clearest case. The companion tool (Evidence Coalition Explorer) is a nice touch - small but genuinely functional, letting readers toggle evidence and see conclusions change themselves. Would be even stronger with a second person double-checking the classifications.
Two Witnesses presents a thoughtful way to analyse organisational AI incidents at the level of individual factual claims rather than treating the incident as a single narrative. The key contribution is asking which stakeholder’s evidence is actually necessary to establish each claim and testing that question by removing developer evidence, outside evidence, or the connection between the two.
The “then versus now” distinction is particularly valuable. The attribution example shows that a claim may initially require evidence held by more than one organization even though a later disclosure eventually makes one source sufficient on its own. This highlights an important evidence preservation problem that the records and identifiers needed to connect organizations’ timelines may themselves be safety relevant evidence.
The fine grained 134 piece analysis is also a useful robustness check because it tests whether apparent multi-stakeholder dependence is simply an artifact of grouping several facts into one claim. The result appropriately narrows the strongest finding to a small number of cases rather than overstating the prevalence of coalition dependent evidence.
The most important next step is independent replication. A preregistered coding protocol, multiple independent coders, and a broader incident set spanning multiple developers would help determine whether the observed pattern generalizes beyond these two cases. Overall, this is a useful and carefully scoped analytical method with a promising application to attribution, disclosure, and cross-organizational incident reconstruction.
Read full reviewShow less
Cite this project
@misc{thuo2026two,
title = {{Two Witnesses: An evidentiary coalition audit of AI-agent incident disclosure}},
author = {Michelle Wanjiku Thuo},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/two-witnesses-an-evidentiary-coalition-audit-of-aiagent-incident-disclosure-oxcm}},
url = {https://apartresearch.com/sprints/projects/two-witnesses-an-evidentiary-coalition-audit-of-aiagent-incident-disclosure-oxcm}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …