Modeling the political process to forecast the outcomes of hypothetical AI governance proposals
Linh Le, David Williams-King · Team LIDA
Submitted to The AI Forecasting Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
International cooperation on AI governance faces a fundamental trust problem: countries like the US and China struggle to assess whether proposed agreements would actually be implemented by their counterpart's domestic political systems. This uncertainty undermines the credibility of commitments and hinders the establishment of safety regulations needed to prevent catastrophic AI risks. We address this challenge by developing a system that predicts whether AI safety legislation would gain support within a country's government, enabling both domestic policymakers and international partners to evaluate the political feasibility of proposed regulations. Our approach uses large language models to generate interpretable yes/no questions about legislative bills, then learns legislator-specific perspective representations that capture individual voting patterns on AI policy. We collect and analyze voting records from the 118th and 119th U.S. Congresses (2024-2025), identifying 146 AI safety-related bills. Our model significantly outperforms baseline approaches in forecasting senatorial votes on AI legislation. Additionally, we develop a suite of hypothetical AI governance policies ranging from strict to permissive, using our model to identify political feasibility thresholds—the boundaries between policies likely to pass versus fail. This work provides a concrete tool for improving transparency and trust in international AI governance negotiations.
Reviews
A novel and an exciting approach - I’d be excited to see you work on this more!
However, I’m quite confused by how the accuracy of the model is measured. The way I understand this works now is:
-When training, the model learns from the sponsorship patterns
-In testing, you ask whether a senator that is mentioned in the bill would support it. This feels odd to me - doesn’t the fact that a senator is mentioned in the bill mean that they support the bill?
It would be beneficial to hold out entire bills from training and then test accuracy on those, so that you can measure accuracy more accurately.
Currently, there’s no validation for the main use case. You generate predictions for the gradated hypothetical policies, but for them, there is no ground truth.
I also think the baseline comparison is flawed - GPT3-oss is not a state-of-the-art AI forecasting tool (you could have compared to this, for example: https://safe.ai/blog/forecasting). And the cutoff date makes the comparison quite unfair.
Read full reviewShow less
* While it seems to me that the actual contribution is quite interesting, the framing leaves a lot to be desired. The data used is from the 118th and 119th US Congress, but the framing is about international actors measuring feasibility of agreements being followed by specific nations. While the US Congress may approve a bill if it has been introduced from one of its own, buy in for international regulations is a meaningfully different scenario which doesn't map perfectly onto data used. Nonetheless, I think the contribution is quite interesting and could be leveraged to understand what self-imposed national level legislation is possible.
* The general idea described in the "Question Representation" section is interesting, but I am having trouble picking up on what exact information is provided to the model.
* Why filter to only the bills which mention "risks"? It would be my intuition that the other bills may have provided additional insight into how certain members of Congress feel about AI. Why was this decision made (e.g. cost saving)?
* Would like to see a more principled way of developing the tested bills, although this is definitely sufficient for a hackathon.
* Table 1 is confusing... what is actually being presented here? Why would you show the different versions of the policy and their variants compared to GPT4-oss-20B if it wasn't actually used to evaluate said policies?
* I would have been quite curious to see a more thorough exploration of how this differs from prior policymaker vote forecasting works. While it is certainly possible that the approaches introduced here build meaningfully on existing literature, it is hard to tell if this is the case.
* I would have liked to see more detail and focus on the novel contributions of this work, which I primarily see as the LLM-based question/answer framework used, and how one might improve this to make it more accurate.
* Lastly, I'd be curious what the naive approach of simply assuming that a policymaker would vote for the bill if it was introduced by a majority from their own party. My guess is that this baseline would definitely be above 0%, so it feels necessary to include.
Read full reviewShow less
Cite this project
@misc{le2025modeling,
title = {{Modeling the political process to forecast the outcomes of hypothetical AI governance proposals}},
author = {Linh Le and David Williams-King},
year = {2025},
month = nov,
note = {Submitted to The AI Forecasting Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/modeling-the-political-process-to-forecast-the-outcomes-of-hypothetical-ai-governance-proposals-0n2c}},
url = {https://apartresearch.com/sprints/projects/modeling-the-political-process-to-forecast-the-outcomes-of-hypothetical-ai-governance-proposals-0n2c}
}More from The AI Forecasting Hackathon
- View project: System Dynamics Game-Theoretic Model of the AI Development Race
System Dynamics Game-Theoretic Model of the AI Development Race
System Dynamics BCN
A Game theoretic / System Dynamics model of the race dynamics of the US, China, and EU, as a follow up to the Armstrong et al. (2016) paper “Racing to the Precipice”. We find preliminary results where knowledge of …
- View project: ExogenousAI
ExogenousAI
Fibonacci
Current AI capability forecasting methodologies, including EpochAI's Direct Approach and Biological Anchors framework, primarily rely on internal metrics such as training compute and scaling laws while assuming stable …
- View project: AI Incidents Forecasting
AI Incidents Forecasting
KLACE
This research develops a framework for forecasting AI incidents to help predict future risks. We have developed two models that forecasts incidents which include calibrated 90% prediction intervals with backtests. These …