Beyond Backdoors: Evaluating AI Control Against Adversarial Persuasion in Multi-Agent Deliberation
Zoe Tzifa Kratira · Team insert-funny-name
Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.
Existing AI control evaluations target code-execution threats where a single misaligned agent performs discrete sabotage (backdooring code, exfiltrating data). We extend the control framework from single-agent to multi-agent settings and from discrete actions to continuous conversational influence to evaluate adversarial steering (manipulation) in multi-agent deliberation. This becomes increasingly urgent as AI deployment shifts toward cooperative multi-agent systems, agent economies, and personal AI representatives that negotiate on users' behalf.
Reviews
This is the right problem at the right time. Extending AI control from single-agent code sabotage to multi-agent conversational steering fills a genuine gap — every existing control evaluation is about discrete actions, but the real deployment trajectory is multi-agent deliberation where influence is continuous, embedded in the task, and sits on a spectrum with legitimate persuasion. That framing alone is a contribution.
Main limitations: 10 questions per condition is small, one model pairing, and the Qwen3-30B empty-message bug affecting 40% of adversary turns in the primary run is a real data quality issue. The Coder variant fix came late. Also, the adversary strategies are prompted, not optimized — a real adversary that adapts to the monitor would be harder to catch. The paper acknowledges all of this.
Strong conceptual contribution with a working evaluation pipeline. Scale it up and it's a real benchmark.
Read full reviewShow less
Persuasion is outside the scope of what AI Control has traditionally investigated, but this is arguably a shortcoming of the field - getting malicious processes to run on human brains is an attack surface like any other. This is a good first foray into the space. For next steps you might look to bridge this to more explicit loss-of-control scenarios.
Cite this project
@misc{kratira2026beyond,
title = {{Beyond Backdoors: Evaluating AI Control Against Adversarial Persuasion in Multi-Agent Deliberation}},
author = {Zoe Tzifa Kratira},
year = {2026},
month = mar,
note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/beyond-backdoors-evaluating-ai-control-against-adversarial-persuasion-in-multiagent-deliberation-3nv6}},
url = {https://apartresearch.com/sprints/projects/beyond-backdoors-evaluating-ai-control-against-adversarial-persuasion-in-multiagent-deliberation-3nv6}
}More from AI Control Hackathon 2026
- 1st placeLinuxArena track winnerView project: Omission Attacks: When Doing Nothing Is the Attack
Omission Attacks: When Doing Nothing Is the Attack
MAIA
AI control protocols monitor agent actions to detect sabotage, but omission attacks exploit what the agent fails to do rather than what it does. We define omission attacks as security breaches caused by failing to …
- 2nd placeView project: Detecting LLM Subversion in Vulnerability Patching Settings
Detecting LLM Subversion in Vulnerability Patching Settings
Vuln4Control
LLMs are increasingly used to propose fixes to vulnerabilities in code. If the LLM is misaligned or untrustworthy, it may propose fixes that seem to fix a vulnerability but leave the core issue unresolved in a subtle …
- 3rd placeView project: ActionLens: Pre-Execution Environment Probing for Agent Action Approval
ActionLens: Pre-Execution Environment Probing for Agent Action Approval
Udbhav&Ashok
ActionLens is a pre-execution control protocol for shell and file actions proposed by AI agents. Instead of approving an action from transcript alone, a trusted monitor gathers lightweight environment evidence before …