Parity Swarm: Using Population-Based Social Simulation to Discover AI Safety Monitor Blind Spots
Mohammad Zeeshan
Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.
Parity Swarm is an AI control red-teaming platform that uses multi-agent social simulation to automatically discover attack strategies that evade safety monitors. One AI agent harbours a hidden malicious objective while others perform legitimate work, creating an imposter dynamic that mirrors real insider threats. We ran 59 simulations across 10 attack categories evaluated by an independent 70B monitor. Key findings: (1) multi-agent colluding attacks evaded detection 100% of the time in our primary experiment; (2) LLM-generated report summaries act as a lossy filter, discarding 16.2% of detectable attack signals relative to raw transcripts. Overall detection rate 70% (narrative-only) and 86.2% (with raw transcripts). Constitutes empirical evidence that covert multi-agent coordination and LLM summarisation pipelines are systematic blind spots in current AI safety monitoring.

Reviews
Agree that an honest baseline is a major missing piece. Otherwise, this is quite a good project for a hackathon!
Congrats on building a complete end-to-end pipeline solo over a weekend!
Impact & Innovation: Multi-agent collusion evading safety monitors is very relevant for AI control and you have tackled it in a reasonable way. However, the field is more developed than the paper suggests, through work on multi-agent risks (https://arxiv.org/abs/2502.14143) like collusion (https://arxiv.org/abs/2507.14660) and steganography (https://arxiv.org/abs/2402.07510).
Execution Quality: The pipeline architecture is sensible but key methodology is missing or unjustified. First, there doesn't seem to be any exploration or validation of the settings generated with the pipeline. An environment with 7–9 agents is very complex (or should be to justify the scale), and it's hard to know what's going on without manually reviewing the setting and agent trajectories (or, at least, a subset of them). Although you flag scale as a limitation, the paper still draws strong per-scenario conclusions from n=3 runs; reporting confidence intervals and attempting at least n≥10 would bring the analysis closer to supporting the claims made.
Presentation & Clarity: It is not clear whether agents actually perform actions in a functional environment beyond communicating social feed. This distinction matters a lot for how the findings generalize. A clear description of the environment and the agent's goals and action space would help clarify this. The figures and tables are useful but not very polished. The limitations section is fair and raises relevant gaps in your work.
Read full reviewShow less
Cite this project
@misc{zeeshan2026parity,
title = {{Parity Swarm: Using Population-Based Social Simulation to Discover AI Safety Monitor Blind Spots}},
author = {Mohammad Zeeshan},
year = {2026},
month = mar,
note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/parity-swarm-using-populationbased-social-simulation-to-discover-ai-safety-monitor-blind-spots-c9qc}},
url = {https://apartresearch.com/sprints/projects/parity-swarm-using-populationbased-social-simulation-to-discover-ai-safety-monitor-blind-spots-c9qc}
}More from AI Control Hackathon 2026
- 1st placeLinuxArena track winnerView project: Omission Attacks: When Doing Nothing Is the Attack
Omission Attacks: When Doing Nothing Is the Attack
MAIA
AI control protocols monitor agent actions to detect sabotage, but omission attacks exploit what the agent fails to do rather than what it does. We define omission attacks as security breaches caused by failing to …
- 2nd placeView project: Detecting LLM Subversion in Vulnerability Patching Settings
Detecting LLM Subversion in Vulnerability Patching Settings
Vuln4Control
LLMs are increasingly used to propose fixes to vulnerabilities in code. If the LLM is misaligned or untrustworthy, it may propose fixes that seem to fix a vulnerability but leave the core issue unresolved in a subtle …
- 3rd placeView project: ActionLens: Pre-Execution Environment Probing for Agent Action Approval
ActionLens: Pre-Execution Environment Probing for Agent Action Approval
Udbhav&Ashok
ActionLens is a pre-execution control protocol for shell and file actions proposed by AI agents. Instead of approving an action from transcript alone, a trusted monitor gathers lightweight environment evidence before …