Skip to content
Sprint projectNov 3, 2025Buenos Aires

Table Top Agents

Luca De Leo · Team Just Luca

Submitted to The AI Forecasting Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

We present Tabletop Agents, an AI-powered framework that accelerates AI governance scenario exploration by orchestrating autonomous AI agents through structured tabletop exercises. Traditional policy wargaming takes years to iterate—RAND's cycles span 4 years for 43 exercises. AI capabilities advance faster than policy preparation can accommodate, creating a critical tempo mismatch. Tabletop Agents compresses preparation cycles from years to minutes while maintaining strategic fidelity. Our working prototype successfully orchestrates multi-agent, multi-turn scenarios where autonomous agents communicate via CLI, persist state in SQLite, and coordinate through turn-based phases. A 2-agent, 2-turn test scenario executed in 4:40 with 5 messages exchanged, demonstrating core orchestration mechanics. The framework enables researchers to run dozens of scenario variations per week instead of months between exercises, generating empirical data on AI governance strategic dynamics at scale. By automating the role-playing that traditionally requires extensive human coordination, we provide better data, faster iteration, and realistic practice during the critical pre-AGI window.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

Does the project meaningfully advance AI timeline prediction and capability forecasting? Does it clearly connect to measurable indicators of AI progress (compute, benchmarks, economic impacts, automation milestones)? Does it build on or challenge existing forecasting frameworks like biological anchors, scaling laws, or scenario planning? Does it offer novel methodologies, data sources, or empirical insights that could improve forecast accuracy? Is it grounded in observable trends rather than pure speculation?

Does this project inform critical decisions about AI development and preparedness? Does it help identify key uncertainties, decision points, or early warning indicators? How well does the project connect technical metrics to real-world impacts and policy needs? Could the output guide resource allocation, safety research priorities, or regulatory timelines? Does it reduce uncertainty around transformative AI milestones or capability emergence?

Is the project methodologically rigorous, reproducible, and technically sound? Is the forecasting approach well-calibrated with appropriate uncertainty quantification? Are the data sources, assumptions, and limitations clearly documented? Does the project demonstrate sound statistical methodology and honest treatment of model uncertainties? Would the tool, model, or framework be useful for ongoing forecasting efforts, research planning, or policy analysis?

  1. The project is about setting up table-top exercises that can be played through I agents, which can speed up policy wargaming. The anecdote that RAND takes 4+ years to do a single scenario is a compelling reason for this project to be done.

    I generally think this looks like a promising approach, but I think there were some key details missing. For example, what were the scenarios tested?

    In general, I would also like to see some thought on why these agents would be useful, and whether they would have meaningful behavior that are comparable to human experts. Would they mirror humans well, and so the outcome of these simulations be useful in themselves? Even if they were as accurate, could they be somewhat useful to humans doing forecasting as reference?

    I would also encourage the authors to look at _why_ RAND processes take so slowly: Which parts of the processes are the bottleneck? And is there any part of the bottleneck that is tractable to solve in this multi-agent scenario? Trying to replace the entire process seems difficult to justify in terms of results.

    Read full reviewShow less
  2. * Contributions are clearly stated and motivated, although a reference to the wargaming cycles noted would be valuable for corroboration of underlying claim

    * Immediate issue: how can we trust that this is actually similar to how experts and/or governments would behave in analogous situations? For example, while it's great that runs are seeded and therefore can be replicated, but how would these agents behavior change if e.g. the prompt contained the same content but was ordered differently?

    * This reads as LLM written, which isn't disqualifying in itself, but I would like to see more citations and reference to specific previous works

    * It seems to me that the actual scenario run is quite important to the value of this submission; how was this defined? Do we know that this is representative of actual wargame instances? How might we show this is the case?

    * I do think this is an interesting proposal, and could be a good contribution if the above issues are resolved, i.e. (i) sensitivity analysis which demonstrates arbitrary/immaterial modifications, e.g. to the prompting of agents, do not meaningfully change outcomes in a statistically significant manner, (ii) thoroughly demonstrating that a given scenario is designed in line with traditional wargaming approaches, and (iii) running multiple baseline comparisons where a traditional wargame and TTA results are compared to each other.

    Read full reviewShow less

Cite this project

@misc{leo2025table,
  title = {{Table Top Agents}},
  author = {Luca De Leo},
  year = {2025},
  month = nov,
  note = {Submitted to The AI Forecasting Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/table-top-agents-i2zx}},
  url = {https://apartresearch.com/sprints/projects/table-top-agents-i2zx}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026