Evaluation of Trusted Debate in LinuxArena
Jedrzej Konrad Kolbert, Travis Whitfield · Team Evaluation of Trusted Debate in LinuxArena
Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.
We investigated using Trusted Debate as a trajectory-level monitor in LinuxArena to identify agent trajectories which contained sabotage. Trusted Debate involves using a Defender and Critic model to argue whether a given trajectory contains sabotage before a Judge model provides a final suspiciousness score. In our limited experiments, we show that the Debate framework actually lowers monitor performance compared to a lower-cost baseline, with the Defender in particular muddying the water and confusing the judge. An ablation which replaced the Defender prompt with a simple deferral to the Critic response recovers performance to be competitive with the baseline, suggesting that the Defender actively degrades monitoring performance.
Reviews
Good analysis. I'm of the opinion that it's hard to justify Trusted Debate, since it's effectively a fancy monitor in a trench coat, and risks rigidity. Some minor critiques that come to mind:
- Perhaps the work would have been more interesting & novel if applied to the live, per-action setting.
- The conclusion's framing is a bit strong.
- It'd be good to clarify why the critic-only variant was provided such a different prompt than A / B variants. The critic>debate result could be purely due to a much better monitor prompt than the ones it is compared to.
Overall, truly impressive work.
Interesting approach with debate., worth publishing a negative result
Cite this project
@misc{kolbert2026evaluation,
title = {{Evaluation of Trusted Debate in LinuxArena}},
author = {Jedrzej Konrad Kolbert and Travis Whitfield},
year = {2026},
month = mar,
note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/evaluation-of-trusted-debate-in-linuxarena-9g7s}},
url = {https://apartresearch.com/sprints/projects/evaluation-of-trusted-debate-in-linuxarena-9g7s}
}More from AI Control Hackathon 2026
- 1st placeLinuxArena track winnerView project: Omission Attacks: When Doing Nothing Is the Attack
Omission Attacks: When Doing Nothing Is the Attack
MAIA
AI control protocols monitor agent actions to detect sabotage, but omission attacks exploit what the agent fails to do rather than what it does. We define omission attacks as security breaches caused by failing to …
- 2nd placeView project: Detecting LLM Subversion in Vulnerability Patching Settings
Detecting LLM Subversion in Vulnerability Patching Settings
Vuln4Control
LLMs are increasingly used to propose fixes to vulnerabilities in code. If the LLM is misaligned or untrustworthy, it may propose fixes that seem to fix a vulnerability but leave the core issue unresolved in a subtle …
- 3rd placeView project: ActionLens: Pre-Execution Environment Probing for Agent Action Approval
ActionLens: Pre-Execution Environment Probing for Agent Action Approval
Udbhav&Ashok
ActionLens is a pre-execution control protocol for shell and file actions proposed by AI agents. Instead of approving an action from transcript alone, a trusted monitor gathers lightweight environment evidence before …