Skip to content
Sprint projectNov 3, 2025Abuja, Nigeria

Forecasting AGI: A Granular, CHC-Based Approach

Habeeb Abdulfatah · Team Habyb

Submitted to The AI Forecasting Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Forecasting AGI: A Granular, CHC-Based Approach

Share

The paper introduces a data-driven framework for forecasting Artificial General Intelligence (AGI) based on the Cattell-Horn-Carroll (CHC) theory of cognition. It breaks AGI into ten measurable cognitive domains and maps each to existing AI benchmarks. Using GPT-4 (2023) and projected GPT-5 (2025) data, the study applies exponential trend extrapolation to predict human-level proficiency across domains. Results show rapid progress in reading, writing, and math by 2028, but major bottlenecks in memory and reasoning until the 2030s. The approach provides a granular, reproducible, and governance-relevant method for tracking AI progress and informing strategic planning.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

Does the project meaningfully advance AI timeline prediction and capability forecasting? Does it clearly connect to measurable indicators of AI progress (compute, benchmarks, economic impacts, automation milestones)? Does it build on or challenge existing forecasting frameworks like biological anchors, scaling laws, or scenario planning? Does it offer novel methodologies, data sources, or empirical insights that could improve forecast accuracy? Is it grounded in observable trends rather than pure speculation?

Does this project inform critical decisions about AI development and preparedness? Does it help identify key uncertainties, decision points, or early warning indicators? How well does the project connect technical metrics to real-world impacts and policy needs? Could the output guide resource allocation, safety research priorities, or regulatory timelines? Does it reduce uncertainty around transformative AI milestones or capability emergence?

Is the project methodologically rigorous, reproducible, and technically sound? Is the forecasting approach well-calibrated with appropriate uncertainty quantification? Are the data sources, assumptions, and limitations clearly documented? Does the project demonstrate sound statistical methodology and honest treatment of model uncertainties? Would the tool, model, or framework be useful for ongoing forecasting efforts, research planning, or policy analysis?

  1. Building upon recent work on “A Definition of AGI” is sensible and timely. However, the project would benefit from more qualitative justifications of some of the predictions, and also from using some other forecasting methodology in addition to trend extrapolation. For example, your model predicts that Long-Term Memory Storage will reach 100% by 2038, despite both GPT-4 and GPT-5 scoring zero on that factor. Why?

    Some of the graphs are a bit odd, race to 100% proficiency” - it isn’t aligned with the forecasts listed in the table

  2. The work tries to operationalize CHC domains by mapping them into measurable benchmarks. They then try to forecast when AI would hit the 100% proficiency level for each domain.

    I think building it off CHC is a good idea, and it is somewhat fair to weigh each category equally as a starting point. However, I'm not sure what you are doing exactly to do the forecasting: How are you doing exponential trend extrapolation given two data points? The forecasting aspect of this work seems to be making big jumps in logic, and more details (and justification) around implementation details would be nice.

    There is also an assumption that the currently existing set of benchmarks is sufficient, and if there exists a model that can solve all reading-related benchmarks that we would deem the model to be human-level at reading. It is often the case that new benchmarks pop up over time measuring a different aspect of a skill, especially if current models are completely hopeless at that aspect, so you may also want to define a % that all benchmarks cover instead of 100%.

    Read full reviewShow less

Cite this project

@misc{abdulfatah2025forecasting,
  title = {{Forecasting AGI: A Granular, CHC-Based Approach}},
  author = {Habeeb Abdulfatah},
  year = {2025},
  month = nov,
  note = {Submitted to The AI Forecasting Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/forecasting-agi-a-granular-chcbased-approach-9n7z}},
  url = {https://apartresearch.com/sprints/projects/forecasting-agi-a-granular-chcbased-approach-9n7z}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026