Forecasting Autonomous AI Bio-Threat Design Capabilities: Six Models Converge on 2031
Ram Potham
Submitted to The AI Forecasting Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
This paper forecasts when frontier AI models will first achieve a critical threshold for autonomous biological threat design capabilities that cause existential risk. I develop a standardized 100-point evaluation framework and use superforecasting methodology with six independent quantitative models to analyze current AI capabilities in protein design, biosecurity screening, and autonomous research systems.
Reviews
The topic is very relevant, and I like that you consider different modes of forecasting. However, at its current state, I have low confidence in the forecasts produced by this project.
Documentation leaves room to be improved. For example, it’s left unclear why Model 2 and Model 4 were excluded. For Model 5, it’s unclear to me what the different experts forecasted. Model 6 doesn’t really do scenario analysis despite its name - it just multiplies some numbers together.
The project could have benefitted from engaging more deeply with existing literature on the topic. Rather than coming up with a “standardized test” yourself, you could have used an existing benchmark, such as the Virology Capabilities Test, and forecasted performance by models on that benchmark. That would’ve been a more feasible project given the time limit.
I think forecasting biological capabilities is an important question! However, I think there are quite a few implementation decisions that were left out that seem somewhat arbitrary. For example, why were Models 2 and 4 excluded from the weighting? How were the weights for the other models chosen? For the Expert Consensus model, are the estimates across the different organisations comparable (ie are they trying to estimate the same thing)? How was the 100-point framework created, why are the four categories weighted equally, and why do you consider 75 points as a critical threshold?
The more complex the approach is, the more decision points and assumptions you will have to justify. You should consider simplifying the whole approach, or reduce the scope of your research question.
Cite this project
@misc{potham2025forecasting,
title = {{Forecasting Autonomous AI Bio-Threat Design Capabilities: Six Models Converge on 2031}},
author = {Ram Potham},
year = {2025},
month = nov,
note = {Submitted to The AI Forecasting Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/forecasting-autonomous-ai-biothreat-design-capabilities-six-models-converge-on-2031-3gpn}},
url = {https://apartresearch.com/sprints/projects/forecasting-autonomous-ai-biothreat-design-capabilities-six-models-converge-on-2031-3gpn}
}More from The AI Forecasting Hackathon
- View project: System Dynamics Game-Theoretic Model of the AI Development Race
System Dynamics Game-Theoretic Model of the AI Development Race
System Dynamics BCN
A Game theoretic / System Dynamics model of the race dynamics of the US, China, and EU, as a follow up to the Armstrong et al. (2016) paper “Racing to the Precipice”. We find preliminary results where knowledge of …
- View project: ExogenousAI
ExogenousAI
Fibonacci
Current AI capability forecasting methodologies, including EpochAI's Direct Approach and Biological Anchors framework, primarily rely on internal metrics such as training compute and scaling laws while assuming stable …
- View project: AI Incidents Forecasting
AI Incidents Forecasting
KLACE
This research develops a framework for forecasting AI incidents to help predict future risks. We have developed two models that forecasts incidents which include calibrated 90% prediction intervals with backtests. These …