Skip to content
Sprint projectNov 2, 2025New York City

AI Capability Terrain

Pranati Modumudi · Team AI Capabilities Terrain

Submitted to The AI Forecasting Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

AI Capability Terrain Map is an interactive 3D early warning system that visualizes AI capabilities as literal terrain — with mountain peaks representing mastered abilities, rising hills showing emerging skills, and red sinkholes exposing tasks that AI systems should handle easily but consistently fail at.

The system integrates real-time benchmark data from Epoch AI to monitor capability progression and forecast future developments across more than 30 distinct AI domains. By combining empirical performance data with logistic growth modeling and Monte Carlo uncertainty quantification, it produces granular, capability-specific forecasts with confidence intervals — moving beyond generic “AGI timelines” to concrete capability trajectories.

Forecasting and Analysis Pipeline The forecasting component models benchmark performance over time using logistic growth curves fitted to historical data. Each forecast includes a 95% confidence interval, generated via Monte Carlo simulation, allowing users to account for uncertainty in future capability trajectories. The system also detects capability sinkholes — areas where progress stalls despite strong performance in related domains — revealing blind spots in current AI architectures and training paradigms.

Visualization and Interaction The 3D interface, built with React and Three.js, translates forecast data into an intuitive spatial landscape. Users can explore capability clusters, rising trends, and sinkholes in real time. Each terrain region dynamically updates as new benchmark results are integrated, making the map both an analytic and exploratory tool for researchers and policymakers.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

Does the project meaningfully advance AI timeline prediction and capability forecasting? Does it clearly connect to measurable indicators of AI progress (compute, benchmarks, economic impacts, automation milestones)? Does it build on or challenge existing forecasting frameworks like biological anchors, scaling laws, or scenario planning? Does it offer novel methodologies, data sources, or empirical insights that could improve forecast accuracy? Is it grounded in observable trends rather than pure speculation?

Does this project inform critical decisions about AI development and preparedness? Does it help identify key uncertainties, decision points, or early warning indicators? How well does the project connect technical metrics to real-world impacts and policy needs? Could the output guide resource allocation, safety research priorities, or regulatory timelines? Does it reduce uncertainty around transformative AI milestones or capability emergence?

Is the project methodologically rigorous, reproducible, and technically sound? Is the forecasting approach well-calibrated with appropriate uncertainty quantification? Are the data sources, assumptions, and limitations clearly documented? Does the project demonstrate sound statistical methodology and honest treatment of model uncertainties? Would the tool, model, or framework be useful for ongoing forecasting efforts, research planning, or policy analysis?

  1. * I appreciate the scale normalization, as this is a consistent thorn in the side of anyone looking at benchmarks on the meta-level.

    * It isn't clear to me how the capability sinkholes were chosen to begin with.

    * I am quite confused by what is being presented in Figure 1. The 95% confidence intervals either read as oddly long, or incredibly narrow. I assume the latter, but due to how it's presented, I don't really know what these mean.

    * Good use of data sufficiency criteria.

    * I'd be curious what would happen if different benchmarks which alleged to measure the same concepts as those used were introduced. Likely we would see a sharp reduction in the capability being measured.

    * Unfortunately, the terrain map is quite difficult to parse, and not super visually appealing. I would lean into the main contributions and de-emphasize the terrain map (unless you already know a way to make it much more accessible).

    * I think this work would strongly benefit from the use of additional data sources to inform the confidence intervals.

    Read full reviewShow less

Cite this project

@misc{modumudi2025ai,
  title = {{AI Capability Terrain}},
  author = {Pranati Modumudi},
  year = {2025},
  month = nov,
  note = {Submitted to The AI Forecasting Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/ai-capability-terrain-ixd9}},
  url = {https://apartresearch.com/sprints/projects/ai-capability-terrain-ixd9}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026