The Half-Life of Compute Thresholds
Aaron Potter
Submitted to The Technical AI Governance Challenge. Sprint projects are early-stage work by participants, not Apart Research publications.
ompute thresholds are widely discussed as a practical governance tool, but the literature rarely specifies how quickly fixed thresholds become outdated as algorithmic efficiency improves. We develop a compute-threshold staleness model that distinguishes raw training compute from baseline-equivalent (“effective”) compute and defines staleness as the ratio between a policy threshold and the time-varying threshold that would remain aligned with a fixed risk cutoff
Reviews
Not clear on counterfactual impact of this work; doubling time only, without these formulae, already provides "a simple model which can provide a good intuitive grasp of the problem for policy-makers." In practice, this model may be making the the explanation more opaque to policymakers.
There could be some value in describing staleness in the intervals between doubling times; e.g how stale is the threshold at 7 months with an 8 month doubling time; but this still can be reasoned about pretty intuitively with knowledge of the doubling time, and a simple, one-off calculation.
I would want to see more exploration or elaboration on specific use-cases for this work to better understand its impact potential. Maybe dealing with uncertainty around doubling time, or enabling very precise update thresholds under certain conditions.
The model's most significant limitation is the assumption that τ is constant across domains and over time. Algorithmic efficiency likely improves at different rates for different capability types (language understanding vs. code generation vs. chemistry), which means a single threshold update cadence may be insufficient — the paper should explore domain-specific τ values and their policy implications. The backtest, while a nice touch, confirms consistency with Ho et al.'s estimates rather than providing independent validation, since the model's dynamics depend on those same estimates.
Cite this project
@misc{potter2026halflife,
title = {{The Half-Life of Compute Thresholds}},
author = {Aaron Potter},
year = {2026},
month = feb,
note = {Submitted to The Technical AI Governance Challenge, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/the-halflife-of-compute-thresholds-260u}},
url = {https://apartresearch.com/sprints/projects/the-halflife-of-compute-thresholds-260u}
}More from The Technical AI Governance Challenge
- 1st placeView project: LidaSim: Testing AI Policies With Persona-Based Simulations
LidaSim: Testing AI Policies With Persona-Based Simulations
Lida Safety
We simulate well-known figures in AI and politics with agents, scraping large amounts of data to get realistic simulations. Then, we test questions and proposed policies against these public figures, to see which …
- 2nd placeView project: Markov Chain Lock Watermarking: Provably Secure Authentication for LLM Outputs
Markov Chain Lock Watermarking: Provably Secure Authentication for LLM Outputs
MCL
We present Markov Chain Lock (MCL) watermarking, a cryptographically secure framework for authenticating LLM outputs. MCL constrains token generation to follow a secret Markov chain over SHA-256 vocabulary partitions. …
- 3rd placeView project: Political Intelligence for AI Safety: The AI Risk Attitudes Survey (AIRAS)
Political Intelligence for AI Safety: The AI Risk Attitudes Survey (AIRAS)
AIRAS
The AI safety and governance community is making progress on defining red lines around existential risk from advanced AI systems, and building verification infrastructure to support this objective. However, this is only …