AI Capability Terrain
Pranati Modumudi · Team AI Capabilities Terrain
Submitted to The AI Forecasting Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
AI Capability Terrain Map is an interactive 3D early warning system that visualizes AI capabilities as literal terrain — with mountain peaks representing mastered abilities, rising hills showing emerging skills, and red sinkholes exposing tasks that AI systems should handle easily but consistently fail at.
The system integrates real-time benchmark data from Epoch AI to monitor capability progression and forecast future developments across more than 30 distinct AI domains. By combining empirical performance data with logistic growth modeling and Monte Carlo uncertainty quantification, it produces granular, capability-specific forecasts with confidence intervals — moving beyond generic “AGI timelines” to concrete capability trajectories.
Forecasting and Analysis Pipeline The forecasting component models benchmark performance over time using logistic growth curves fitted to historical data. Each forecast includes a 95% confidence interval, generated via Monte Carlo simulation, allowing users to account for uncertainty in future capability trajectories. The system also detects capability sinkholes — areas where progress stalls despite strong performance in related domains — revealing blind spots in current AI architectures and training paradigms.
Visualization and Interaction The 3D interface, built with React and Three.js, translates forecast data into an intuitive spatial landscape. Users can explore capability clusters, rising trends, and sinkholes in real time. Each terrain region dynamically updates as new benchmark results are integrated, making the map both an analytic and exploratory tool for researchers and policymakers.
Reviews
* I appreciate the scale normalization, as this is a consistent thorn in the side of anyone looking at benchmarks on the meta-level.
* It isn't clear to me how the capability sinkholes were chosen to begin with.
* I am quite confused by what is being presented in Figure 1. The 95% confidence intervals either read as oddly long, or incredibly narrow. I assume the latter, but due to how it's presented, I don't really know what these mean.
* Good use of data sufficiency criteria.
* I'd be curious what would happen if different benchmarks which alleged to measure the same concepts as those used were introduced. Likely we would see a sharp reduction in the capability being measured.
* Unfortunately, the terrain map is quite difficult to parse, and not super visually appealing. I would lean into the main contributions and de-emphasize the terrain map (unless you already know a way to make it much more accessible).
* I think this work would strongly benefit from the use of additional data sources to inform the confidence intervals.
Read full reviewShow less
Cite this project
@misc{modumudi2025ai,
title = {{AI Capability Terrain}},
author = {Pranati Modumudi},
year = {2025},
month = nov,
note = {Submitted to The AI Forecasting Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/ai-capability-terrain-ixd9}},
url = {https://apartresearch.com/sprints/projects/ai-capability-terrain-ixd9}
}More from The AI Forecasting Hackathon
- View project: System Dynamics Game-Theoretic Model of the AI Development Race
System Dynamics Game-Theoretic Model of the AI Development Race
System Dynamics BCN
A Game theoretic / System Dynamics model of the race dynamics of the US, China, and EU, as a follow up to the Armstrong et al. (2016) paper “Racing to the Precipice”. We find preliminary results where knowledge of …
- View project: ExogenousAI
ExogenousAI
Fibonacci
Current AI capability forecasting methodologies, including EpochAI's Direct Approach and Biological Anchors framework, primarily rely on internal metrics such as training compute and scaling laws while assuming stable …
- View project: AI Incidents Forecasting
AI Incidents Forecasting
KLACE
This research develops a framework for forecasting AI incidents to help predict future risks. We have developed two models that forecasts incidents which include calibrated 90% prediction intervals with backtests. These …