A Comparative Analysis of Ensemble Protocols for AI Control
Ariel Monzon, Natan Katz, Inbar Shulman, Mike Vazan, Or Shemesh, Guy Porat · Team Heron Blue Team
Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.
We empirically evaluate ensembles of AI control protocols and show that performance depends on signal diversity rather than the number of tools: small, diverse ensembles outperform larger, redundant ones. Correlation between tools limits gains from stacking, while consensus-based aggregation improves robustness, providing a principled basis for ensemble design.
Reviews
Congrats on a very thorough project! Mapping ensemble interactions across 100+ tool combinations is ambitious for a hackathon weekend. The report is really clear and well presented.
Impact & Innovation: The research question, although not groundbreaking, is interesting and pulling it off successfully would provide safety gains. The two novel detection mechanisms are well-motivated. As you flagged, the non-adaptive adversary does bound how directly these ensembles transfer to deployment, though a full red team exploration is understandably beyond hackathon scope.
Execution Quality: The experimental design is impressively thorough. Two gaps worth noting:
- The comparison between ensembles and individual tools isn't compute-matched. A 3-tool ensemble gets three inference passes per sample while a single protocol gets one. Therefore, it's unclear whether the gains come from complementary signals or simply from more compute. The right control would be to run each individual protocol multiple times and aggregate those results, then compare against a tool ensemble at equal compute budget. Without this, the central claim about diversity vs. quantity is confounded.
- There are no confidence intervals in the results or error bars in the figures. Adding them would help calibrate the relevance of the results.
Presentation & Clarity: The paper is well-written and concise with a clear narrative arc. Figures are well-chosen and informative. Minor notes: some figure captions carry substantial analytical claims that would read better in the main text.
Read full reviewShow less
Strong empirical contribution! Really enjoyed reading through the project.
Exhaustively evaluating all tool combinations across multiple aggregation methods is thorough and well-designed. The central finding (mechanism diversity matters more than tool count, and the best small ensemble outperforms the full ensemble) is counterintuitive and immediately actionable. The two novel detection mechanisms are both clever. ACP exploiting backdoor brittleness for semantically preserving perturbations is a particularly nice insight.
Cite this project
@misc{monzon2026comparative,
title = {{A Comparative Analysis of Ensemble Protocols for AI Control}},
author = {Ariel Monzon and Natan Katz and Inbar Shulman and Mike Vazan and Or Shemesh and Guy Porat},
year = {2026},
month = mar,
note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/a-comparative-analysis-of-ensemble-protocols-for-ai-control-0usy}},
url = {https://apartresearch.com/sprints/projects/a-comparative-analysis-of-ensemble-protocols-for-ai-control-0usy}
}More from AI Control Hackathon 2026
- 1st placeLinuxArena track winnerView project: Omission Attacks: When Doing Nothing Is the Attack
Omission Attacks: When Doing Nothing Is the Attack
MAIA
AI control protocols monitor agent actions to detect sabotage, but omission attacks exploit what the agent fails to do rather than what it does. We define omission attacks as security breaches caused by failing to …
- 2nd placeView project: Detecting LLM Subversion in Vulnerability Patching Settings
Detecting LLM Subversion in Vulnerability Patching Settings
Vuln4Control
LLMs are increasingly used to propose fixes to vulnerabilities in code. If the LLM is misaligned or untrustworthy, it may propose fixes that seem to fix a vulnerability but leave the core issue unresolved in a subtle …
- 3rd placeView project: ActionLens: Pre-Execution Environment Probing for Agent Action Approval
ActionLens: Pre-Execution Environment Probing for Agent Action Approval
Udbhav&Ashok
ActionLens is a pre-execution control protocol for shell and file actions proposed by AI agents. Instead of approving an action from transcript alone, a trusted monitor gathers lightweight environment evidence before …