Skip to content
Sprint projectNov 2, 2025UK

AI Incidents Forecasting

Chamod Kalupahana, Ahmed Elbashir, Christopher L. Lübbers · Team KLACE

Submitted to The AI Forecasting Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

This research develops a framework for forecasting AI incidents to help predict future risks. We have developed two models that forecasts incidents which include calibrated 90% prediction intervals with backtests. These models predicts a large growth in the total number of incidents over the next five years. This has generated forecasts with quantified uncertainty, enabling policymakers to shift from reactive to proactive risk mitigation, supporting evidence-based regulation and responsible AI deployment.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

Does the project meaningfully advance AI timeline prediction and capability forecasting? Does it clearly connect to measurable indicators of AI progress (compute, benchmarks, economic impacts, automation milestones)? Does it build on or challenge existing forecasting frameworks like biological anchors, scaling laws, or scenario planning? Does it offer novel methodologies, data sources, or empirical insights that could improve forecast accuracy? Is it grounded in observable trends rather than pure speculation?

Does this project inform critical decisions about AI development and preparedness? Does it help identify key uncertainties, decision points, or early warning indicators? How well does the project connect technical metrics to real-world impacts and policy needs? Could the output guide resource allocation, safety research priorities, or regulatory timelines? Does it reduce uncertainty around transformative AI milestones or capability emergence?

Is the project methodologically rigorous, reproducible, and technically sound? Is the forecasting approach well-calibrated with appropriate uncertainty quantification? Are the data sources, assumptions, and limitations clearly documented? Does the project demonstrate sound statistical methodology and honest treatment of model uncertainties? Would the tool, model, or framework be useful for ongoing forecasting efforts, research planning, or policy analysis?

  1. * Great call using a preexisting work for data preprocessing!

    * Why did you make the design decisions you did for the model (trend-change hinge @ 2021; spline flexibility constraint)? Make sure all decisions like this are thoroughly justified, and that this is conveyed to the reader.

    * Figure 1 uses a scale which makes it quite difficult to evaluate how well your model does on already logged data. Additionally, I think it is unlikely that incidents will continue on the upward trend indefinitely; as some point, like any technological adoption, there will be an inflection point and the rate of change will reduce. This should be taken into account, or at least mentioned, when presenting this kind of plot.

    * You mention that incidents are distinct from reports, which definitely should be mentioned, but I think this could also be used as another feature to broaden the context provided. Did number of reports have similar behavior to number of incidents? What inferences might we be able to draw from this additional source of information?

    * It would also be interesting if an analysis could be conducted on the number of incidents that happens in a given year vs. the number of incidents reported here. Although we can't know this for certain, one could attempt to approximate this with comparisons to other crowd-sourced reporting endeavors and historic analysis. This is similar to the "future work" you present.

    Read full reviewShow less
  2. I think this line of research is quite exciting, and trying to build models on the AI incidents database sounds like a great idea. I like how you included backtesting.

    One failure mode this may have is that it assumes ceteris paribus, i.e. that all conditions (e.g. safety regulations, regulations and conditions of mandatory reporting, etc.) stay the same over the near future, which is unlikely.

Cite this project

@misc{kalupahana2025ai,
  title = {{AI Incidents Forecasting}},
  author = {Chamod Kalupahana and Ahmed Elbashir and Christopher L. Lübbers},
  year = {2025},
  month = nov,
  note = {Submitted to The AI Forecasting Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/ai-incidents-forecasting-w92p}},
  url = {https://apartresearch.com/sprints/projects/ai-incidents-forecasting-w92p}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026