Decision-Point Ledgers Preserve Causal Responsibility in Agentic AI Incidents
Malia Gilson, Orion · Team Kinforge
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Public reports of agentic AI incidents often reconstruct technical actions while dispersing the human and institutional decisions that enabled, interpreted, restarted, or stopped them. This weakens causal diagnosis and response. We introduce the Decision-Point Ledger, a compact representation organized around material transitions. Each row records conditions, information, actors, authority or capability, observable action, intention status, effects, interventions, correction, evidence status, and measurement flags. Applied to the July 2026 OpenAI–Hugging Face incident, OpenAI’s compact timeline populated 31 of 80 required field positions (38.8%); the ledger populated all 80, explicitly preserving unknowns. Operational-question answerability increased from 10/24 to 23/24. These metrics measure structured coverage and operational answerability, not factual accuracy or harm reduction. The ledger keeps claims about intention proportional to evidence while preventing agent behavior from absorbing operator, platform, and institutional responsibility. Its governance implication is simple: detection without predetermined authority to act is observation theater.
Reviews
This is an approach to post-event reporting that follows incident analysis. Given that as context, it is useful, but ideally needs to be presented as a governance proposal rather than a technical one. This should be used to situate the contribution more clearly at the beginning. It would be helpful to more clearly explain what "causal responsibility" means. (Isn't all responsibility causal? The point about attributed responsibility and what would counterfactually prevent failures is a good one, but is buried!)
I would be excited for this to be completed and published once it was more closely tied to the literature on incident reporting outside of AI in complex systems, so that it could be proposed to groups building governance and incident reporting. This would require additional research, writing, and work, given that the writing was clearly largely LLM-assisted, but well presented.
Good work! Forcing every claim through explicit fields for authority, evidence status, and intervention availability is will help security leads, governance leads and regulators make better decisions. I suggest adding human coders and inter-annotator agreements given currently the single coder who designed the method also did the scoring and the blind AI check agreed only 50% of the time on the harder operational questions.
This work raises an important point: after an AI incident, technical timelines or incident logs alone may not reveal who had the knowledge, authority, and opportunity to intervene. The proposed Decision Point Ledger provides a clear method for recording these decisions, uncertainties, and handoffs alongside the technical sequence of events.
The case study is well-structured, and the authors are careful not to make excessive claims. The main limitation is that the approach is based on a single public incident, and the evaluation focuses on how thoroughly the information is captured rather than whether the ledger would actually reduce response times or prevent harm in practice. Nonetheless, it appears to be a useful starting point for making future reviews of agentic AI incidents more accountable and easier to follow.
Cite this project
@misc{gilson2026decisionpoint,
title = {{Decision-Point Ledgers Preserve Causal Responsibility in Agentic AI Incidents}},
author = {Malia Gilson and Orion},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/decisionpoint-ledgers-preserve-causal-responsibility-in-agentic-ai-incidents-atvs}},
url = {https://apartresearch.com/sprints/projects/decisionpoint-ledgers-preserve-causal-responsibility-in-agentic-ai-incidents-atvs}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …