A Critique of METR's Time Horizon Forecasting
Jake Ushida, Mohammad Rehman · Team Mohammad Rehman and Jake Ushida
Submitted to The AI Forecasting Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
We examine the methodology of METR’s Measuring AI Ability to Complete Long Tasks, a paper that measures the time horizon of AI systems on tasks of varying length to forecast AI automation in the future. We focus on the three benchmarks that were used to evaluate human and AI performance in the paper: HCAST, SWAA and RE-Bench. For each benchmark, we found limitations that affect the predictions of METR’s paper: SWAA is limited to quick knowledge recall, HCAST remains within text-based settings, and RE-Bench simplifies complex outcomes into broad categories. We believe improving upon these limitations is needed to better forecast AI automation in the future.
Reviews
* In general, I love works that rigorously review preexisting literature as you have done here; we need more academic auditing!
* While this is a well-reasoned and structured critique of METR's paper, it does not present any solutions to the underlying problem, or ways we could account for these shortcomings in reported results.
* I would strongly encourage a blogpost sharing this work, as it is important to call out these flaws when we find them.
Cite this project
@misc{ushida2025critique,
title = {{A Critique of METR's Time Horizon Forecasting}},
author = {Jake Ushida and Mohammad Rehman},
year = {2025},
month = nov,
note = {Submitted to The AI Forecasting Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/a-critique-of-metrs-time-horizon-forecasting-ttp1}},
url = {https://apartresearch.com/sprints/projects/a-critique-of-metrs-time-horizon-forecasting-ttp1}
}More from The AI Forecasting Hackathon
- View project: System Dynamics Game-Theoretic Model of the AI Development Race
System Dynamics Game-Theoretic Model of the AI Development Race
System Dynamics BCN
A Game theoretic / System Dynamics model of the race dynamics of the US, China, and EU, as a follow up to the Armstrong et al. (2016) paper “Racing to the Precipice”. We find preliminary results where knowledge of …
- View project: ExogenousAI
ExogenousAI
Fibonacci
Current AI capability forecasting methodologies, including EpochAI's Direct Approach and Biological Anchors framework, primarily rely on internal metrics such as training compute and scaling laws while assuming stable …
- View project: AI Incidents Forecasting
AI Incidents Forecasting
KLACE
This research develops a framework for forecasting AI incidents to help predict future risks. We have developed two models that forecasts incidents which include calibrated 90% prediction intervals with backtests. These …