Table Top Agents
Luca De Leo · Team Just Luca
Submitted to The AI Forecasting Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
We present Tabletop Agents, an AI-powered framework that accelerates AI governance scenario exploration by orchestrating autonomous AI agents through structured tabletop exercises. Traditional policy wargaming takes years to iterate—RAND's cycles span 4 years for 43 exercises. AI capabilities advance faster than policy preparation can accommodate, creating a critical tempo mismatch. Tabletop Agents compresses preparation cycles from years to minutes while maintaining strategic fidelity. Our working prototype successfully orchestrates multi-agent, multi-turn scenarios where autonomous agents communicate via CLI, persist state in SQLite, and coordinate through turn-based phases. A 2-agent, 2-turn test scenario executed in 4:40 with 5 messages exchanged, demonstrating core orchestration mechanics. The framework enables researchers to run dozens of scenario variations per week instead of months between exercises, generating empirical data on AI governance strategic dynamics at scale. By automating the role-playing that traditionally requires extensive human coordination, we provide better data, faster iteration, and realistic practice during the critical pre-AGI window.
Reviews
The project is about setting up table-top exercises that can be played through I agents, which can speed up policy wargaming. The anecdote that RAND takes 4+ years to do a single scenario is a compelling reason for this project to be done.
I generally think this looks like a promising approach, but I think there were some key details missing. For example, what were the scenarios tested?
In general, I would also like to see some thought on why these agents would be useful, and whether they would have meaningful behavior that are comparable to human experts. Would they mirror humans well, and so the outcome of these simulations be useful in themselves? Even if they were as accurate, could they be somewhat useful to humans doing forecasting as reference?
I would also encourage the authors to look at _why_ RAND processes take so slowly: Which parts of the processes are the bottleneck? And is there any part of the bottleneck that is tractable to solve in this multi-agent scenario? Trying to replace the entire process seems difficult to justify in terms of results.
Read full reviewShow less
* Contributions are clearly stated and motivated, although a reference to the wargaming cycles noted would be valuable for corroboration of underlying claim
* Immediate issue: how can we trust that this is actually similar to how experts and/or governments would behave in analogous situations? For example, while it's great that runs are seeded and therefore can be replicated, but how would these agents behavior change if e.g. the prompt contained the same content but was ordered differently?
* This reads as LLM written, which isn't disqualifying in itself, but I would like to see more citations and reference to specific previous works
* It seems to me that the actual scenario run is quite important to the value of this submission; how was this defined? Do we know that this is representative of actual wargame instances? How might we show this is the case?
* I do think this is an interesting proposal, and could be a good contribution if the above issues are resolved, i.e. (i) sensitivity analysis which demonstrates arbitrary/immaterial modifications, e.g. to the prompting of agents, do not meaningfully change outcomes in a statistically significant manner, (ii) thoroughly demonstrating that a given scenario is designed in line with traditional wargaming approaches, and (iii) running multiple baseline comparisons where a traditional wargame and TTA results are compared to each other.
Read full reviewShow less
Cite this project
@misc{leo2025table,
title = {{Table Top Agents}},
author = {Luca De Leo},
year = {2025},
month = nov,
note = {Submitted to The AI Forecasting Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/table-top-agents-i2zx}},
url = {https://apartresearch.com/sprints/projects/table-top-agents-i2zx}
}More from The AI Forecasting Hackathon
- View project: System Dynamics Game-Theoretic Model of the AI Development Race
System Dynamics Game-Theoretic Model of the AI Development Race
System Dynamics BCN
A Game theoretic / System Dynamics model of the race dynamics of the US, China, and EU, as a follow up to the Armstrong et al. (2016) paper “Racing to the Precipice”. We find preliminary results where knowledge of …
- View project: ExogenousAI
ExogenousAI
Fibonacci
Current AI capability forecasting methodologies, including EpochAI's Direct Approach and Biological Anchors framework, primarily rely on internal metrics such as training compute and scaling laws while assuming stable …
- View project: AI Incidents Forecasting
AI Incidents Forecasting
KLACE
This research develops a framework for forecasting AI incidents to help predict future risks. We have developed two models that forecasts incidents which include calibrated 90% prediction intervals with backtests. These …