AI Incident Timeline & Evidence Dashboard
Hüseyin Tenlik, Emre Bilgiç, Deniz Süren · Team 10
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
The 2026 OpenAI–Hugging Face incident produced overlapping but non-identical public accounts from the organizations involved, independent investigators, and secondary analysts. We ask whether a claim-level evidence model can make agreement, disagreement, and uncertainty easier to verify. We reviewed six incident-specific public sources and encoded 30 dated events and interpretive claims, each linked to a primary source and labeled Confirmed, Disputed, or Uncertain under an explicit rubric. We then built a React/Express dashboard that supports filtering, record-level evidence inspection, source links, and eight evidence-derived “Checks / Watch Next” questions. The final dataset contains 21 Confirmed, five Disputed, and four Uncertain records. The main result is not a new forensic finding, but a transparent evidence layer that preserves source conflicts and open questions rather than collapsing the incident into one narrative.
Reviews
This project is well executed but I struggle to see how useful it is in the grander scheme of AI security. I could see future versions of this being implemented in existing AI security databases as a way to track ongoing incidents.
The strongest part of this project is the decision to treat the incident as a set of evidence rather than trying to force everything into one clean story. That works especially well where the sources disagree. The stopping-date issue is a good example: one source effectively ends the campaign on 12 July, while Hugging Face still records 1,130 actions on 13 July. The restart timeline has a similar problem, and both would be easy to miss in a normal narrative. The JSON dataset, schema, API filtering, citation files, MIT licence and CI workflow make this something other researchers could genuinely reuse, not just a dashboard. The main weakness is source coverage, especially the missing 37-page technical report released with the August post, which may affect the 27 June alert classification. I would also separate first-party confirmation from independent corroboration, since 11 of the 21 Confirmed records rely only on first-party reporting. Disputed should also distinguish actual source conflicts from interpretations being tested. Finally, the checks would be much stronger if each had a clear pass, fail and resolution condition.
Read full reviewShow less
Cite this project
@misc{tenlik2026ai,
title = {{AI Incident Timeline \& Evidence Dashboard}},
author = {Hüseyin Tenlik and Emre Bilgiç and Deniz Süren},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/ai-incident-timeline-evidence-dashboard-vfjq}},
url = {https://apartresearch.com/sprints/projects/ai-incident-timeline-evidence-dashboard-vfjq}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …