Skip to content
Sprint projectNov 2, 2025Montreal, Canada

Modeling the political process to forecast the outcomes of hypothetical AI governance proposals

Linh Le, David Williams-King · Team LIDA

Submitted to The AI Forecasting Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Modeling the political process to forecast the outcomes of hypothetical AI governance proposals

Share

International cooperation on AI governance faces a fundamental trust problem: countries like the US and China struggle to assess whether proposed agreements would actually be implemented by their counterpart's domestic political systems. This uncertainty undermines the credibility of commitments and hinders the establishment of safety regulations needed to prevent catastrophic AI risks. We address this challenge by developing a system that predicts whether AI safety legislation would gain support within a country's government, enabling both domestic policymakers and international partners to evaluate the political feasibility of proposed regulations. Our approach uses large language models to generate interpretable yes/no questions about legislative bills, then learns legislator-specific perspective representations that capture individual voting patterns on AI policy. We collect and analyze voting records from the 118th and 119th U.S. Congresses (2024-2025), identifying 146 AI safety-related bills. Our model significantly outperforms baseline approaches in forecasting senatorial votes on AI legislation. Additionally, we develop a suite of hypothetical AI governance policies ranging from strict to permissive, using our model to identify political feasibility thresholds—the boundaries between policies likely to pass versus fail. This work provides a concrete tool for improving transparency and trust in international AI governance negotiations.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

Does the project meaningfully advance AI timeline prediction and capability forecasting? Does it clearly connect to measurable indicators of AI progress (compute, benchmarks, economic impacts, automation milestones)? Does it build on or challenge existing forecasting frameworks like biological anchors, scaling laws, or scenario planning? Does it offer novel methodologies, data sources, or empirical insights that could improve forecast accuracy? Is it grounded in observable trends rather than pure speculation?

Does this project inform critical decisions about AI development and preparedness? Does it help identify key uncertainties, decision points, or early warning indicators? How well does the project connect technical metrics to real-world impacts and policy needs? Could the output guide resource allocation, safety research priorities, or regulatory timelines? Does it reduce uncertainty around transformative AI milestones or capability emergence?

Is the project methodologically rigorous, reproducible, and technically sound? Is the forecasting approach well-calibrated with appropriate uncertainty quantification? Are the data sources, assumptions, and limitations clearly documented? Does the project demonstrate sound statistical methodology and honest treatment of model uncertainties? Would the tool, model, or framework be useful for ongoing forecasting efforts, research planning, or policy analysis?

  1. A novel and an exciting approach - I’d be excited to see you work on this more!

    However, I’m quite confused by how the accuracy of the model is measured. The way I understand this works now is:

    -When training, the model learns from the sponsorship patterns

    -In testing, you ask whether a senator that is mentioned in the bill would support it. This feels odd to me - doesn’t the fact that a senator is mentioned in the bill mean that they support the bill?

    It would be beneficial to hold out entire bills from training and then test accuracy on those, so that you can measure accuracy more accurately.

    Currently, there’s no validation for the main use case. You generate predictions for the gradated hypothetical policies, but for them, there is no ground truth.

    I also think the baseline comparison is flawed - GPT3-oss is not a state-of-the-art AI forecasting tool (you could have compared to this, for example: https://safe.ai/blog/forecasting). And the cutoff date makes the comparison quite unfair.

    Read full reviewShow less
  2. * While it seems to me that the actual contribution is quite interesting, the framing leaves a lot to be desired. The data used is from the 118th and 119th US Congress, but the framing is about international actors measuring feasibility of agreements being followed by specific nations. While the US Congress may approve a bill if it has been introduced from one of its own, buy in for international regulations is a meaningfully different scenario which doesn't map perfectly onto data used. Nonetheless, I think the contribution is quite interesting and could be leveraged to understand what self-imposed national level legislation is possible.

    * The general idea described in the "Question Representation" section is interesting, but I am having trouble picking up on what exact information is provided to the model.

    * Why filter to only the bills which mention "risks"? It would be my intuition that the other bills may have provided additional insight into how certain members of Congress feel about AI. Why was this decision made (e.g. cost saving)?

    * Would like to see a more principled way of developing the tested bills, although this is definitely sufficient for a hackathon.

    * Table 1 is confusing... what is actually being presented here? Why would you show the different versions of the policy and their variants compared to GPT4-oss-20B if it wasn't actually used to evaluate said policies?

    * I would have been quite curious to see a more thorough exploration of how this differs from prior policymaker vote forecasting works. While it is certainly possible that the approaches introduced here build meaningfully on existing literature, it is hard to tell if this is the case.

    * I would have liked to see more detail and focus on the novel contributions of this work, which I primarily see as the LLM-based question/answer framework used, and how one might improve this to make it more accurate.

    * Lastly, I'd be curious what the naive approach of simply assuming that a policymaker would vote for the bill if it was introduced by a majority from their own party. My guess is that this baseline would definitely be above 0%, so it feels necessary to include.

    Read full reviewShow less

Cite this project

@misc{le2025modeling,
  title = {{Modeling the political process to forecast the outcomes of hypothetical AI governance proposals}},
  author = {Linh Le and David Williams-King},
  year = {2025},
  month = nov,
  note = {Submitted to The AI Forecasting Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/modeling-the-political-process-to-forecast-the-outcomes-of-hypothetical-ai-governance-proposals-0n2c}},
  url = {https://apartresearch.com/sprints/projects/modeling-the-political-process-to-forecast-the-outcomes-of-hypothetical-ai-governance-proposals-0n2c}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026