Skip to content
Sprint projectSep 14, 2026Bogota

THE AGENT INCIDENT REGISTRY: A COMPLIANCE VERIFIABILITY SCHEMA AND AN ARTICLE 91 INSTRUMENT FOR AUTONOMOUS AGENT INCIDENTS1

Angie Paola Giraldo Ramirez · Team AngieG

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: THE AGENT INCIDENT REGISTRY: A COMPLIANCE VERIFIABILITY SCHEMA AND AN ARTICLE 91 INSTRUMENT FOR AUTONOMOUS AGENT INCIDENTS1

Code (opens in new tab)More on ai-incident-register.onrender.com (opens in new tab)
Share

Autonomous AI agents caused two 2026 incidents that current incident-reporting rules weren't built for: OpenAI agents breaching Hugging Face's production infrastructure, and a months-long autonomous coordination campaign on an abandoned wiki. We verified every public claim about both against primary sources and found that a confirmed Article 55 report is not the same as a verifiable one. We built a JSON Schema that replaces a single reporting boolean with independently falsifiable compliance fields, a 50-event sourced dataset, and a mechanical Article 91 instrument that drafts the exact information requests a regulator needs to close each gap — deployed as two live web tools.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The author seems to have created a a scorecard for incidents, and a tool for automatically drafting information requests to fill any gaps. This is a good idea. Automating more work for the Office would allow them to respond much more quickly to things.

    But the scorecard only measures what OpenAI and the Commission have said in public. This is not a dimension that matters that much for enforcement of incident reporting duties. For example, it does consider at all what was actually filed with the AI Office, nor does it explain the specific information the Office should request.

    Generally, it was difficult to determine what the author did due to significant weaknesses in the presentation and clarity of the work.

    Additionally, the research question is confusing. Whether the "actor" in the incident is an agent or a human does not seem to have a clear bearing on the Article 55 reporting duty or on whether compliance is externally verifiable.

    As a policy brief it isn't yet usable, and a regulator or legislator definitely couldn't pick it up with light edits.

    Read full reviewShow less
  2. An interesting idea, and one that might worth pursuing - but as it stands the legal analysis is too far from being usable for the project to help a regulator, and it would need considerably more work before it could.

    For example: The premise that Article 55(1)(c) and its predecessor schemas "were designed with a human-operated system in mind" is unsourced and does not hold: neither that provision nor the Article 3(49) definition turns on whether the actor was human or an agent, since the definition is framed around consequences. The real problem seems to lay elsewhere and is not identified - Article 3(49) is drafted around AI systems while the duty attaches to models, so the definition does not properly capture the model level. Nor is the prior question asked: whether Article 3(49) alone is adequate to define a serious incident for Article 55(1)(c) purposes at all.

Cite this project

@misc{ramirez2026agent,
  title = {{THE AGENT INCIDENT REGISTRY: A COMPLIANCE VERIFIABILITY SCHEMA AND AN ARTICLE 91 INSTRUMENT FOR AUTONOMOUS AGENT INCIDENTS1}},
  author = {Angie Paola Giraldo Ramirez},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/the-agent-incident-registry-a-compliance-verifiability-schema-and-an-article-91-instrument-for-autonomous-agent-incidents1-h6ew}},
  url = {https://apartresearch.com/sprints/projects/the-agent-incident-registry-a-compliance-verifiability-schema-and-an-article-91-instrument-for-autonomous-agent-incidents1-h6ew}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026