Skip to content
Sprint projectNov 3, 2025Chicago

Forecasting Autonomous AI Bio-Threat Design Capabilities: Six Models Converge on 2031

Ram Potham

Submitted to The AI Forecasting Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Forecasting Autonomous AI Bio-Threat Design Capabilities: Six Models Converge on 2031

Code (opens in new tab)
Share

This paper forecasts when frontier AI models will first achieve a critical threshold for autonomous biological threat design capabilities that cause existential risk. I develop a standardized 100-point evaluation framework and use superforecasting methodology with six independent quantitative models to analyze current AI capabilities in protein design, biosecurity screening, and autonomous research systems.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

Does the project meaningfully advance AI timeline prediction and capability forecasting? Does it clearly connect to measurable indicators of AI progress (compute, benchmarks, economic impacts, automation milestones)? Does it build on or challenge existing forecasting frameworks like biological anchors, scaling laws, or scenario planning? Does it offer novel methodologies, data sources, or empirical insights that could improve forecast accuracy? Is it grounded in observable trends rather than pure speculation?

Does this project inform critical decisions about AI development and preparedness? Does it help identify key uncertainties, decision points, or early warning indicators? How well does the project connect technical metrics to real-world impacts and policy needs? Could the output guide resource allocation, safety research priorities, or regulatory timelines? Does it reduce uncertainty around transformative AI milestones or capability emergence?

Is the project methodologically rigorous, reproducible, and technically sound? Is the forecasting approach well-calibrated with appropriate uncertainty quantification? Are the data sources, assumptions, and limitations clearly documented? Does the project demonstrate sound statistical methodology and honest treatment of model uncertainties? Would the tool, model, or framework be useful for ongoing forecasting efforts, research planning, or policy analysis?

  1. The topic is very relevant, and I like that you consider different modes of forecasting. However, at its current state, I have low confidence in the forecasts produced by this project.

    Documentation leaves room to be improved. For example, it’s left unclear why Model 2 and Model 4 were excluded. For Model 5, it’s unclear to me what the different experts forecasted. Model 6 doesn’t really do scenario analysis despite its name - it just multiplies some numbers together.

    The project could have benefitted from engaging more deeply with existing literature on the topic. Rather than coming up with a “standardized test” yourself, you could have used an existing benchmark, such as the Virology Capabilities Test, and forecasted performance by models on that benchmark. That would’ve been a more feasible project given the time limit.

  2. I think forecasting biological capabilities is an important question! However, I think there are quite a few implementation decisions that were left out that seem somewhat arbitrary. For example, why were Models 2 and 4 excluded from the weighting? How were the weights for the other models chosen? For the Expert Consensus model, are the estimates across the different organisations comparable (ie are they trying to estimate the same thing)? How was the 100-point framework created, why are the four categories weighted equally, and why do you consider 75 points as a critical threshold?

    The more complex the approach is, the more decision points and assumptions you will have to justify. You should consider simplifying the whole approach, or reduce the scope of your research question.

Cite this project

@misc{potham2025forecasting,
  title = {{Forecasting Autonomous AI Bio-Threat Design Capabilities: Six Models Converge on 2031}},
  author = {Ram Potham},
  year = {2025},
  month = nov,
  note = {Submitted to The AI Forecasting Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/forecasting-autonomous-ai-biothreat-design-capabilities-six-models-converge-on-2031-3gpn}},
  url = {https://apartresearch.com/sprints/projects/forecasting-autonomous-ai-biothreat-design-capabilities-six-models-converge-on-2031-3gpn}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026