Monitor Capture: A Myerson-Satterthwaite Impossibility for AI Oversight Protocols
Publius Dirac, Oscar Fuentes, Christian Vuye, Anantha Shakthi Ganeshan Thevar · Team Monitor Capture
Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.
LLM-based monitors that learn from observed agent behavior can be shaped by the agents they oversee, overfitting to seen attacks while missing novel ones. We call this *monitor capture* and argue it represents a fourth axis of the AI control tradeoff, grounded in the Myerson-Satterthwaite (1983) impossibility result. In our multi-agent marketplace simulation, an adversarial LLM agent covertly moves value past an in-context learning monitor over 10 rounds. A held-out attack set that never enters the monitor's training pool measures the "capture gap" between apparent and true safety. Across 8 seeds (Haiku 4.5 monitor, Sonnet 4.6 adversary), apparent safety reaches 0.58 while true safety stalls at 0.38 — a capture gap of 0.20 — and the adversary moves 237 units of value undetected.
Reviews
The threat model of monitors that learn from observed behaviours being shaped by the agents they oversee is important, and this project executed it solidly with the marketplace setup. The capture gap across all 8 seeds is a useful empirical finding.
This work can be strengthen by testing against static monitors or monitors with different update strategies would help establish how general the capture phenomenon is. I'd also be excited to see this extended to stronger model pairings and with explicit defense mechanisms listed in the limitations (pool resets, diversity-aware example selection, ensemble pairings, etc.) to map out when capture can and can't be mitigated.
Great project. I'd be interested to see how these results generalize to a less toy setting.
Cite this project
@misc{dirac2026monitor,
title = {{Monitor Capture: A Myerson-Satterthwaite Impossibility for AI Oversight Protocols}},
author = {Publius Dirac and Oscar Fuentes and Christian Vuye and Anantha Shakthi Ganeshan Thevar},
year = {2026},
month = mar,
note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/monitor-capture-a-myersonsatterthwaite-impossibility-for-ai-oversight-protocols-qtr5}},
url = {https://apartresearch.com/sprints/projects/monitor-capture-a-myersonsatterthwaite-impossibility-for-ai-oversight-protocols-qtr5}
}More from AI Control Hackathon 2026
- 1st placeLinuxArena track winnerView project: Omission Attacks: When Doing Nothing Is the Attack
Omission Attacks: When Doing Nothing Is the Attack
MAIA
AI control protocols monitor agent actions to detect sabotage, but omission attacks exploit what the agent fails to do rather than what it does. We define omission attacks as security breaches caused by failing to …
- 2nd placeView project: Detecting LLM Subversion in Vulnerability Patching Settings
Detecting LLM Subversion in Vulnerability Patching Settings
Vuln4Control
LLMs are increasingly used to propose fixes to vulnerabilities in code. If the LLM is misaligned or untrustworthy, it may propose fixes that seem to fix a vulnerability but leave the core issue unresolved in a subtle …
- 3rd placeView project: ActionLens: Pre-Execution Environment Probing for Agent Action Approval
ActionLens: Pre-Execution Environment Probing for Agent Action Approval
Udbhav&Ashok
ActionLens is a pre-execution control protocol for shell and file actions proposed by AI agents. Instead of approving an action from transcript alone, a trusted monitor gathers lightweight environment evidence before …