Skip to content
Sprint projectSep 14, 2026Kolkata

MAIR: The Misaligned AI Incident Reporting Standard

Rudrani Ghosh · Team Misaligned AI Incident Reporting Standard

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

When an autonomous agent breaks out of an environment, frameworks like CVSS usually flag it as zero severity. CVSS expects a buffer overflow or an unpatched vulnerability, not an agent abusing valid API credentials or tool permissions. We built MAIR (Misaligned AI Incident Reporting) to address that blind spot. It is a five-axis scoring model and dashboard built specifically for agent containment failures. Instead of relying on qualitative postmortems, MAIR maps narrative incident reports into quantitative scores based on factors like intent ambiguity, escalation depth, and blast radius. The system also maps incident metrics directly against reporting criteria for the EU AI Act and California SB 53, so compliance teams know immediately if a legal threshold was breached. To support practical triage, the platform includes a NetworkX graph of agent propagation paths and a phase by phase defense matrix tested against 30 real world incidents, including the Hugging Face sandbox escape and Anthropic eval breakdowns.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This is an ambitious effort that seeks to provide a complete solution for incident reporting as part of a very rapid sprint; the fact that it falls short in various ways is a function of overambitious goals, and the project itself shows promise.

    The literature review and related work is unfortunately very incomplete, partly because so many of the other related projects are still actively under development, and semi-public. (e.g. ISO AWI 25870 is a non-public working draft; https://www.iso.org/standard/91804.html , and the NIST workshop has evidently not yet led to a publication: https://www.nist.gov/news-events/events/2026/05/nist-workshop-ai-incident-management .) However, this means that many of the otherwise novel contributions here are already being discussed within groups discussing the issue.

    The contribution also proposes a new set of dimensions that are not validated, understandably given the sprint length, and the rater variance used sensitivity instead of more standard and more rigorous inter-rater reliability measures. Proposing that this is the right way forward moves to far, as does saying that this "fixes this problem" on the basis of a single set of events. The overclaiming is actively unhelpful, and undermines the very useful contribution.

    In summary, the work is excellent, but the ambition of claiming this as a new standard is unfortunate, as it would be far more useful and impactful to position it as an alternative metric useful for reporting frameworks being developed.

    Read full reviewShow less
  2. This work proposes MAIR, a framework for rating the severity of AI incidents similarly to CVSS for vulnerabilities.

    The proposal is interesting and clearly presented, but its novelty is unclear given existing frameworks such as the OECD Common Reporting Framework for AI Incidents and OWASP AIVSS.

    More importantly, the paper claims that MAIR can identify when legal reporting requirements are triggered, but does not explain how its scores and thresholds are derived from the cited regulations or provide legal validation of that mapping.

  3. MAIR is a useful attempt to give AI incidents a more structured language than free form post mortems, and the dashboard makes the proposal easy to understand. I like separating egress, detection and intent diff.

    My main concern is validation. Applying the framework to 30 incidents shows that it can be used, but not yet that its ordering or thresholds are correct. I’d next use multiple blinded raters, measure agreement, validate the weights on held-out incidents and keep the legal-reporting mapping separate from the severity score itself.

Cite this project

@misc{ghosh2026mair,
  title = {{MAIR: The Misaligned AI Incident Reporting Standard}},
  author = {Rudrani Ghosh},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/mair-the-misaligned-ai-incident-reporting-standard-juvx}},
  url = {https://apartresearch.com/sprints/projects/mair-the-misaligned-ai-incident-reporting-standard-juvx}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026