Schelling coordination via agentic loops
Olli Järviniemi, Meeri Kuoppala, Callum-Luis Kindred · Team Focal
Submitted to AI Control Hackathon 2025. Sprint projects are early-stage work by participants, not Apart Research publications.
Capability to collude undermines untrusted monitoring protocols. We propose a novel way to evaluate acausal coordination capabilities of LLMs via repeated sampling in causally separated agentic loops. We provide an implementation and run evaluations on Claude 3.7 Sonnet on coordinating on Wikipedia articles. We find that while Sonnet is able to enumerate dozens of Schelling coordination strategies, it is inconsistent on its selection of strategy and performs poorly in more adversarial settings. Our framework is flexible, supporting various modifications to the experimental setup, and we propose several directions for further research. Implementation can be found at https://github.com/ollijarviniemi/Schelling_loop/.
Reviews
No public critique yet.
Cite this project
@misc{jarviniemi2025schelling,
title = {{Schelling coordination via agentic loops}},
author = {Olli Järviniemi and Meeri Kuoppala and Callum-Luis Kindred},
year = {2025},
month = mar,
note = {Submitted to AI Control Hackathon 2025, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/schelling-coordination-via-agentic-loops-azvd}},
url = {https://apartresearch.com/sprints/projects/schelling-coordination-via-agentic-loops-azvd}
}More from AI Control Hackathon 2025
- 1st place by peer reviewView project: Can models use their Chain-of-Thought to attack overseers?
Can models use their Chain-of-Thought to attack overseers?
This project explores vulnerabilities in AI evaluation mechanisms, specifically focusing on how AI agents might influence their overseeing AI judges through manipulative instructions embedded within their …
- 3rd place by peer reviewView project: Model Models: Simulating a Trusted Monitor
Model Models: Simulating a Trusted Monitor
Charlie's Angels
We offer initial investigations into whether the untrusted model can 'simulate' the trusted monitor: is U able to successfully guess what suspicion score T will assign in the APPS setting? We also offer a clean, modular …
- 4th place by peer reviewView project: Adding Document Summaries to Control Arena
Adding Document Summaries to Control Arena
One of the current focuses of AI Control research is finding new settings within which to examine the control problem. In this hackathon, we looked at defining a new AI control setting based around summarizing documents …