Skip to content
Sprint projectNov 2, 2025Barcelona + Istanbul

AI Treaty Momentum Index (ATMI)

Oksana Kotelnikova, Sofiia Lobanova · Team AI Treaty Momentum Index

Submitted to The AI Forecasting Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

We present the AI Treaty Momentum Index (ATMI), a lightweight forecasting model that estimates the near-term chance of a new international AI risk treaty or binding regulation. ATMI learns from past treaties in climate, nuclear, biosafety, and finance by encoding each case with mechanism classes—buckets of factors that help or hinder agreement (e.g., Policy window, Capability shock, Verification feasibility, Public salience, Epistemic shift, Geopolitical alignment). Each mechanism is scored from −5 (hurts) to +5 (helps) per case, and we map these mechanism profiles to outcomes (treaty vs. no treaty). We use the fit in two modes: Nowcasting—apply today’s AI governance signals to produce a current chance with 80% intervals; Scenario-casting—stress-testing futures where scenarios fork from AI-2027 milestones to see how the odds shift. Early findings: the largest marginal lifts come from Epistemic shift and Geopolitical alignment, followed by Public salience and Epistemic infrastructure. As of Nov 2025, ATMI estimates a ~0.6% chance of a new treaty, suppressed by misaligned geopolitics, weak verification readiness, and negative epistemic/resource shifts. We propose to follow this MVP and build a public “anti-doomsday” tool: a web dashboard that tracks each mechanism with live indicators, shows the ATMI curve over time, and flags actionable weakest links (the marginal mechanisms whose improvement most raises the chance of agreement).

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

Does the project meaningfully advance AI timeline prediction and capability forecasting? Does it clearly connect to measurable indicators of AI progress (compute, benchmarks, economic impacts, automation milestones)? Does it build on or challenge existing forecasting frameworks like biological anchors, scaling laws, or scenario planning? Does it offer novel methodologies, data sources, or empirical insights that could improve forecast accuracy? Is it grounded in observable trends rather than pure speculation?

Does this project inform critical decisions about AI development and preparedness? Does it help identify key uncertainties, decision points, or early warning indicators? How well does the project connect technical metrics to real-world impacts and policy needs? Could the output guide resource allocation, safety research priorities, or regulatory timelines? Does it reduce uncertainty around transformative AI milestones or capability emergence?

Is the project methodologically rigorous, reproducible, and technically sound? Is the forecasting approach well-calibrated with appropriate uncertainty quantification? Are the data sources, assumptions, and limitations clearly documented? Does the project demonstrate sound statistical methodology and honest treatment of model uncertainties? Would the tool, model, or framework be useful for ongoing forecasting efforts, research planning, or policy analysis?

  1. * It isn't clear to me what was taken into account to determine the prior medians; given that this is the foundation of your contribution, I'd like to see more justification and explanation of the approach. For example, a case study for one of these would have been quite valuable to see.

    * Another way to connect this further to AI cases would have been to examine specific instances in the recent past where any number of these influences occurred, and what the impact on policy was, if anything. Immediately I can think of the cluster of incidents surrounding deepfake porn, which did catalyze some US legislation.

    * I feel that the "nowcast" is where the bulk of value is introduced, at least from a policymaker perspective. Given that, I would've liked to see more focus on the inputs used here, possibly even including uncertainty analysis.

  2. I like the topic and the conclusions that come out of the project: That there are 'high-leverage' mechanisms that need to happen before an AI treaty can happen.

    On the whole, I think the forecasting aspect of the project is weaker: there are quite a few assumptions (e.g. prior medians, prior range) that sort of are made by the authors that automatically lead to the conclusions as stated. On the whole this is a useful setting for a thought experiment but I would encourage the authors to focus on more realistically modelling some specific criteria (e.g. just on epistemic shift + geopolitical alignment)

Cite this project

@misc{kotelnikova2025ai,
  title = {{AI Treaty Momentum Index (ATMI)}},
  author = {Oksana Kotelnikova and Sofiia Lobanova},
  year = {2025},
  month = nov,
  note = {Submitted to The AI Forecasting Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/ai-treaty-momentum-index-atmi-lst4}},
  url = {https://apartresearch.com/sprints/projects/ai-treaty-momentum-index-atmi-lst4}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026