Bugs over Backdoors: Efficient and Evasive Sabotage via Plausible Code Bugs in LinuxArena
Timothy Obiso, Dean Cahill · Team Plausible Deniability
Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.
AI control attack policies vary widely in efficacy and computational cost. We design and evaluate a bug-based attack policy for LinuxArena, wherein an untrusted agent introduces plausible code defects rather than obfuscated backdoors as its primary sabotage vector. Across nine samples, bug-attack achieves an 89% side-task success rate, exceeding the 78% of the standard obfuscation baseline, while using 14.5x fewer tokens. Notably, bug-attack and standard attack exhibit opposite vulnerability profiles under trusted-editing defense: trusted editing catches 37% of bug-attack sabotages while catching none of standard attack's, suggesting the two are complementary rather than substitutes. We also report single-sample results for two exploratory policies. We treat all findings as preliminary given sample size constraints, but consider the efficiency and evasion results directionally robust.
Reviews
Bug-looking attacks is an interesting red-team strategy to study in the context of control. The results suggest they have higher suspicious rates and are edited away more often than hidden backdoors, but can still be successful attacks with a much lower generation budget. It would be interesting to validate these results over more samples and conduct a comprehensive comparison of the suspicion score distributions and qualitative differences with hidden backdoors attacks.
The paper investigates the performance of agents prompted to make their attacks look like bugs.
Their agents achieved slightly better performance with a slightly better suspicion profile at vastly fewer tokens spent. The downside of the attack is that trusted editing removes many more of their attacks (37%) than with the default policy (none).
In a practical setting, a malicious agent prefers to be perceived as less competent (writing buggy code) than nefarious (writing backdoors), so this approach is relevant to the setting.
I'd have liked to have seen more details on the vast performance improvement (the bug-attack spends fewer tokens than the honest policy).
All in all, this is a relevant exploration in the framework, but not breaking any new ground.
Cite this project
@misc{obiso2026bugs,
title = {{Bugs over Backdoors: Efficient and Evasive Sabotage via Plausible Code Bugs in LinuxArena}},
author = {Timothy Obiso and Dean Cahill},
year = {2026},
month = mar,
note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/bugs-over-backdoors-efficient-and-evasive-sabotage-via-plausible-code-bugs-in-linuxarena-fcp6}},
url = {https://apartresearch.com/sprints/projects/bugs-over-backdoors-efficient-and-evasive-sabotage-via-plausible-code-bugs-in-linuxarena-fcp6}
}More from AI Control Hackathon 2026
- 1st placeLinuxArena track winnerView project: Omission Attacks: When Doing Nothing Is the Attack
Omission Attacks: When Doing Nothing Is the Attack
MAIA
AI control protocols monitor agent actions to detect sabotage, but omission attacks exploit what the agent fails to do rather than what it does. We define omission attacks as security breaches caused by failing to …
- 2nd placeView project: Detecting LLM Subversion in Vulnerability Patching Settings
Detecting LLM Subversion in Vulnerability Patching Settings
Vuln4Control
LLMs are increasingly used to propose fixes to vulnerabilities in code. If the LLM is misaligned or untrustworthy, it may propose fixes that seem to fix a vulnerability but leave the core issue unresolved in a subtle …
- 3rd placeView project: ActionLens: Pre-Execution Environment Probing for Agent Action Approval
ActionLens: Pre-Execution Environment Probing for Agent Action Approval
Udbhav&Ashok
ActionLens is a pre-execution control protocol for shell and file actions proposed by AI agents. Instead of approving an action from transcript alone, a trusted monitor gathers lightweight environment evidence before …