Skip to content
Sprint projectSep 13, 2026Eureka, South Dakota

Decision-Point Ledgers Preserve Causal Responsibility in Agentic AI Incidents

Malia Gilson, Orion · Team Kinforge

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Decision-Point Ledgers Preserve Causal Responsibility in Agentic AI Incidents

Share

Public reports of agentic AI incidents often reconstruct technical actions while dispersing the human and institutional decisions that enabled, interpreted, restarted, or stopped them. This weakens causal diagnosis and response. We introduce the Decision-Point Ledger, a compact representation organized around material transitions. Each row records conditions, information, actors, authority or capability, observable action, intention status, effects, interventions, correction, evidence status, and measurement flags. Applied to the July 2026 OpenAI–Hugging Face incident, OpenAI’s compact timeline populated 31 of 80 required field positions (38.8%); the ledger populated all 80, explicitly preserving unknowns. Operational-question answerability increased from 10/24 to 23/24. These metrics measure structured coverage and operational answerability, not factual accuracy or harm reduction. The ledger keeps claims about intention proportional to evidence while preventing agent behavior from absorbing operator, platform, and institutional responsibility. Its governance implication is simple: detection without predetermined authority to act is observation theater.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This is an approach to post-event reporting that follows incident analysis. Given that as context, it is useful, but ideally needs to be presented as a governance proposal rather than a technical one. This should be used to situate the contribution more clearly at the beginning. It would be helpful to more clearly explain what "causal responsibility" means. (Isn't all responsibility causal? The point about attributed responsibility and what would counterfactually prevent failures is a good one, but is buried!)

    I would be excited for this to be completed and published once it was more closely tied to the literature on incident reporting outside of AI in complex systems, so that it could be proposed to groups building governance and incident reporting. This would require additional research, writing, and work, given that the writing was clearly largely LLM-assisted, but well presented.

  2. Good work! Forcing every claim through explicit fields for authority, evidence status, and intervention availability is will help security leads, governance leads and regulators make better decisions. I suggest adding human coders and inter-annotator agreements given currently the single coder who designed the method also did the scoring and the blind AI check agreed only 50% of the time on the harder operational questions.

  3. This work raises an important point: after an AI incident, technical timelines or incident logs alone may not reveal who had the knowledge, authority, and opportunity to intervene. The proposed Decision Point Ledger provides a clear method for recording these decisions, uncertainties, and handoffs alongside the technical sequence of events.

    The case study is well-structured, and the authors are careful not to make excessive claims. The main limitation is that the approach is based on a single public incident, and the evaluation focuses on how thoroughly the information is captured rather than whether the ledger would actually reduce response times or prevent harm in practice. Nonetheless, it appears to be a useful starting point for making future reviews of agentic AI incidents more accountable and easier to follow.

Cite this project

@misc{gilson2026decisionpoint,
  title = {{Decision-Point Ledgers Preserve Causal Responsibility in Agentic AI Incidents}},
  author = {Malia Gilson and Orion},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/decisionpoint-ledgers-preserve-causal-responsibility-in-agentic-ai-incidents-atvs}},
  url = {https://apartresearch.com/sprints/projects/decisionpoint-ledgers-preserve-causal-responsibility-in-agentic-ai-incidents-atvs}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026