AI Control through Majority Voting
Andreas Moe, Willam Wale, Joakim Hauger Sunde · Team The Majority
Submitted to AI Control Hackathon 2025. Sprint projects are early-stage work by participants, not Apart Research publications.
This report introduces an AI control protocol designed to safely use powerful but potentially untrustworthy AI for programming tasks. The core problem is that such AI might deliberately insert backdoors into the generated code. The proposed solution queries the untrustworthy AI multiple times (c) with the same problem, but slightly varied prompts (using random tokens). This yields c different solutions. A trusted AI then creates a final program that runs all c solutions and outputs the result determined by a majority vote.
Reviews
I like this alot. I think that adding random tokens to the end of the prompt isn't good enough though. The untrusted model might decide to attack on any input that has the word "prime" in it so adding random token at the end wouldn't change anything. But you coudl try summerizing or changing the prompt in other ways. Overall I think this was a good project though
Cite this project
@misc{moe2025ai,
title = {{AI Control through Majority Voting}},
author = {Andreas Moe and Willam Wale and Joakim Hauger Sunde},
year = {2025},
month = mar,
note = {Submitted to AI Control Hackathon 2025, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/ai-control-through-majority-voting-uiaa}},
url = {https://apartresearch.com/sprints/projects/ai-control-through-majority-voting-uiaa}
}More from AI Control Hackathon 2025
- 1st place by peer reviewView project: Can models use their Chain-of-Thought to attack overseers?
Can models use their Chain-of-Thought to attack overseers?
This project explores vulnerabilities in AI evaluation mechanisms, specifically focusing on how AI agents might influence their overseeing AI judges through manipulative instructions embedded within their …
- 2nd place by peer reviewView project: Schelling coordination via agentic loops
Schelling coordination via agentic loops
Focal
Capability to collude undermines untrusted monitoring protocols. We propose a novel way to evaluate acausal coordination capabilities of LLMs via repeated sampling in causally separated agentic loops. We provide an …
- 3rd place by peer reviewView project: Model Models: Simulating a Trusted Monitor
Model Models: Simulating a Trusted Monitor
Charlie's Angels
We offer initial investigations into whether the untrusted model can 'simulate' the trusted monitor: is U able to successfully guess what suspicion score T will assign in the APPS setting? We also offer a clean, modular …