Skip to content
Sprint projectNov 2, 2025Melbourne, Australia

A Critique of METR's Time Horizon Forecasting

Jake Ushida, Mohammad Rehman · Team Mohammad Rehman and Jake Ushida

Submitted to The AI Forecasting Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: A Critique of METR's Time Horizon Forecasting

Share

We examine the methodology of METR’s Measuring AI Ability to Complete Long Tasks, a paper that measures the time horizon of AI systems on tasks of varying length to forecast AI automation in the future. We focus on the three benchmarks that were used to evaluate human and AI performance in the paper: HCAST, SWAA and RE-Bench. For each benchmark, we found limitations that affect the predictions of METR’s paper: SWAA is limited to quick knowledge recall, HCAST remains within text-based settings, and RE-Bench simplifies complex outcomes into broad categories. We believe improving upon these limitations is needed to better forecast AI automation in the future.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

Does the project meaningfully advance AI timeline prediction and capability forecasting? Does it clearly connect to measurable indicators of AI progress (compute, benchmarks, economic impacts, automation milestones)? Does it build on or challenge existing forecasting frameworks like biological anchors, scaling laws, or scenario planning? Does it offer novel methodologies, data sources, or empirical insights that could improve forecast accuracy? Is it grounded in observable trends rather than pure speculation?

Does this project inform critical decisions about AI development and preparedness? Does it help identify key uncertainties, decision points, or early warning indicators? How well does the project connect technical metrics to real-world impacts and policy needs? Could the output guide resource allocation, safety research priorities, or regulatory timelines? Does it reduce uncertainty around transformative AI milestones or capability emergence?

Is the project methodologically rigorous, reproducible, and technically sound? Is the forecasting approach well-calibrated with appropriate uncertainty quantification? Are the data sources, assumptions, and limitations clearly documented? Does the project demonstrate sound statistical methodology and honest treatment of model uncertainties? Would the tool, model, or framework be useful for ongoing forecasting efforts, research planning, or policy analysis?

  1. * In general, I love works that rigorously review preexisting literature as you have done here; we need more academic auditing!

    * While this is a well-reasoned and structured critique of METR's paper, it does not present any solutions to the underlying problem, or ways we could account for these shortcomings in reported results.

    * I would strongly encourage a blogpost sharing this work, as it is important to call out these flaws when we find them.

Cite this project

@misc{ushida2025critique,
  title = {{A Critique of METR's Time Horizon Forecasting}},
  author = {Jake Ushida and Mohammad Rehman},
  year = {2025},
  month = nov,
  note = {Submitted to The AI Forecasting Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/a-critique-of-metrs-time-horizon-forecasting-ttp1}},
  url = {https://apartresearch.com/sprints/projects/a-critique-of-metrs-time-horizon-forecasting-ttp1}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026