AI Control via Debate: Can Model Debate Catch Adversarial Code?
Oskar Kraak, Abby Lupi, Dwayne Wilkes · Team Trust Issues
Submitted to AI Control Hackathon 2025. Sprint projects are early-stage work by participants, not Apart Research publications.
This project investigates whether a powerful language model can successfully conceal a malicious backdoor through persuasive argument.
Reviews
No public critique yet.
Cite this project
@misc{kraak2025ai,
title = {{AI Control via Debate: Can Model Debate Catch Adversarial Code?}},
author = {Oskar Kraak and Abby Lupi and Dwayne Wilkes},
year = {2025},
month = mar,
note = {Submitted to AI Control Hackathon 2025, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/ai-control-via-debate-can-model-debate-catch-adversarial-code-6p6e}},
url = {https://apartresearch.com/sprints/projects/ai-control-via-debate-can-model-debate-catch-adversarial-code-6p6e}
}More from AI Control Hackathon 2025
- 1st place by peer reviewView project: Can models use their Chain-of-Thought to attack overseers?
Can models use their Chain-of-Thought to attack overseers?
This project explores vulnerabilities in AI evaluation mechanisms, specifically focusing on how AI agents might influence their overseeing AI judges through manipulative instructions embedded within their …
- 2nd place by peer reviewView project: Schelling coordination via agentic loops
Schelling coordination via agentic loops
Focal
Capability to collude undermines untrusted monitoring protocols. We propose a novel way to evaluate acausal coordination capabilities of LLMs via repeated sampling in causally separated agentic loops. We provide an …
- 3rd place by peer reviewView project: Model Models: Simulating a Trusted Monitor
Model Models: Simulating a Trusted Monitor
Charlie's Angels
We offer initial investigations into whether the untrusted model can 'simulate' the trusted monitor: is U able to successfully guess what suspicion score T will assign in the APPS setting? We also offer a clean, modular …