On Multi-agent swarming
Andrew Liang · Team
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Our contributions are as follows: We investigate whether agents assigned to individual tasks can recognize their peers and use their information without explicit collaboration instructions, and how these behaviors depend on activity-trace exposure. We measure how assertions from an authoritative planner agent affect collective answers when a worker holds contradictory evidence, distinguishing worker pushback from adoption of the planner’s incorrect answer. We test whether controlled perturbations of planner messages and worker evidence can identify planner-sensitive answers and localize influential messages without access to reasoning traces.
Reviews
Most methodologically serious entry. The deterministic replay framework is the standout, 210 byte-identical identity replays so every flip is attributable to its edit. That is a genuinely useful technique. Also honest about non-results. Weaknesses: small samples everywhere (the ROC is ten episodes), one model, and the audit's discriminative power is unresolved, not demonstrated. The conclusion overreaches, 'such a planner should not exist' is not what the data shows. Scale the question set and this is a paper.
Ambitious and interesting work. The two studies carried out in this project tackle a kind of important question after the recent agent swarm incidents, primarily how information and authority move between agents. I specifically liked the planner-influence study and its implementation methods. The main limitations as stated are that all the findings rest on one model and very small screened samples so the deference rate is directional rather than settled, and the shared artifact study is more of illustrative case observations than a measured effect yet. Overall a strong, careful piece of work that would benefit most from running the planner study at the n=150-200 you mention.
Cite this project
@misc{liang2026multiagent,
title = {{On Multi-agent swarming}},
author = {Andrew Liang},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/on-multiagent-swarming-0h4g}},
url = {https://apartresearch.com/sprints/projects/on-multiagent-swarming-0h4g}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …