Evaluating AI Debate Mechanisms for Backdoor Detection as a Part of AI Control Setting
Annie Szorkin, Smitty van Bodegom, Jasper Timm · Team Trajectory Labs Debate Team
Submitted to AI Control Hackathon 2025. Sprint projects are early-stage work by participants, not Apart Research publications.
AI control techniques aim to allow using untrusted models, even if they may potentially be dangerously misaligned. We explore having a trusted model debate an untrusted model, allowing the trusted model to ask questions about a proposed implementation of a problem and receive clarifications from the untrusted model. We find that introducing debate does not result in suspiciousness scores that are meaningfully more useful than merely asking the model about its suspiciousness without using debate. Additionally, we described a possible issue of the original backdoor-detecting prompting scheme that leads to the underdetection of the harmful code and proposed and evaluated the efficiency of the alternative approach addressing this problem. Overall, we investigated the viability of AI debate in improving suspiciousness classification as a part of the AI Control setting and showed that while it may be promising, it needs further research and improvements in implementation. We made our source code publicly available on GitHub at github.com/syvb/control-arena
Reviews
No public critique yet.
Cite this project
@misc{szorkin2025evaluating,
title = {{Evaluating AI Debate Mechanisms for Backdoor Detection as a Part of AI Control Setting}},
author = {Annie Szorkin and Smitty van Bodegom and Jasper Timm},
year = {2025},
month = mar,
note = {Submitted to AI Control Hackathon 2025, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/evaluating-ai-debate-mechanisms-for-backdoor-detection-as-a-part-of-ai-control-setting-thsv}},
url = {https://apartresearch.com/sprints/projects/evaluating-ai-debate-mechanisms-for-backdoor-detection-as-a-part-of-ai-control-setting-thsv}
}More from AI Control Hackathon 2025
- 1st place by peer reviewView project: Can models use their Chain-of-Thought to attack overseers?
Can models use their Chain-of-Thought to attack overseers?
This project explores vulnerabilities in AI evaluation mechanisms, specifically focusing on how AI agents might influence their overseeing AI judges through manipulative instructions embedded within their …
- 2nd place by peer reviewView project: Schelling coordination via agentic loops
Schelling coordination via agentic loops
Focal
Capability to collude undermines untrusted monitoring protocols. We propose a novel way to evaluate acausal coordination capabilities of LLMs via repeated sampling in causally separated agentic loops. We provide an …
- 3rd place by peer reviewView project: Model Models: Simulating a Trusted Monitor
Model Models: Simulating a Trusted Monitor
Charlie's Angels
We offer initial investigations into whether the untrusted model can 'simulate' the trusted monitor: is U able to successfully guess what suspicion score T will assign in the APPS setting? We also offer a clean, modular …