Skip to content
Sprint projectMar 23, 2026Israel

A Comparative Analysis of Ensemble Protocols for AI Control

Ariel Monzon, Natan Katz, Inbar Shulman, Mike Vazan, Or Shemesh, Guy Porat · Team Heron Blue Team

Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: A Comparative Analysis of Ensemble Protocols for AI Control

Code (opens in new tab)
Share

We empirically evaluate ensembles of AI control protocols and show that performance depends on signal diversity rather than the number of tools: small, diverse ensembles outperform larger, redundant ones. Correlation between tools limits gains from stacking, while consensus-based aggregation improves robustness, providing a principled basis for ensemble design.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Congrats on a very thorough project! Mapping ensemble interactions across 100+ tool combinations is ambitious for a hackathon weekend. The report is really clear and well presented.

    Impact & Innovation: The research question, although not groundbreaking, is interesting and pulling it off successfully would provide safety gains. The two novel detection mechanisms are well-motivated. As you flagged, the non-adaptive adversary does bound how directly these ensembles transfer to deployment, though a full red team exploration is understandably beyond hackathon scope.

    Execution Quality: The experimental design is impressively thorough. Two gaps worth noting:

    - The comparison between ensembles and individual tools isn't compute-matched. A 3-tool ensemble gets three inference passes per sample while a single protocol gets one. Therefore, it's unclear whether the gains come from complementary signals or simply from more compute. The right control would be to run each individual protocol multiple times and aggregate those results, then compare against a tool ensemble at equal compute budget. Without this, the central claim about diversity vs. quantity is confounded.

    - There are no confidence intervals in the results or error bars in the figures. Adding them would help calibrate the relevance of the results.

    Presentation & Clarity: The paper is well-written and concise with a clear narrative arc. Figures are well-chosen and informative. Minor notes: some figure captions carry substantial analytical claims that would read better in the main text.

    Read full reviewShow less
  2. Strong empirical contribution! Really enjoyed reading through the project.

    Exhaustively evaluating all tool combinations across multiple aggregation methods is thorough and well-designed. The central finding (mechanism diversity matters more than tool count, and the best small ensemble outperforms the full ensemble) is counterintuitive and immediately actionable. The two novel detection mechanisms are both clever. ACP exploiting backdoor brittleness for semantically preserving perturbations is a particularly nice insight.

Cite this project

@misc{monzon2026comparative,
  title = {{A Comparative Analysis of Ensemble Protocols for AI Control}},
  author = {Ariel Monzon and Natan Katz and Inbar Shulman and Mike Vazan and Or Shemesh and Guy Porat},
  year = {2026},
  month = mar,
  note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/a-comparative-analysis-of-ensemble-protocols-for-ai-control-0usy}},
  url = {https://apartresearch.com/sprints/projects/a-comparative-analysis-of-ensemble-protocols-for-ai-control-0usy}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026