Sabotage Under the Ratchet: Can Autonomous Research Agents Hide Backdoors While Optimizing Metrics?
Achintya Paningapalli, Arthur Wigo · Team wTeam
Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.
We investigate whether autonomous AI research agents can optimize a target metric while covertly inserting harmful behavior. Using Karpathy’s autoresearch framework, we model a dual-objective agent that must both improve validation performance (val_bpb) and embed a trigger-based backdoor. We introduce two attack strategies—data poisoning and architectural modification—and a principled trigger selection method based on large-scale token frequency analysis. To evaluate defenses, we design a monitoring framework comparing stateless and stateful approaches across both heuristic and LLM-based methods. Our results show that temporally distributed attacks, spread across multiple commits, evade stateless monitoring while being detectable through cumulative analysis. These findings suggest that single-metric oversight is insufficient for autonomous agents and highlight the need for stateful auditing in AI control systems.
Reviews
- The most important next step is completing the GPU training runs. Without actual backdoor insertion and validation, the contribution remains a framework rather than an empirical finding.
- Consider using a stronger agent model than DeepSeek-Coder:33b - the paper itself notes this may underestimate attack sophistication.
Decent idea, needs more work to prove it actually works with longer training and smarter models being used. Good presentation and executive for limited time availability.
Cite this project
@misc{paningapalli2026sabotage,
title = {{Sabotage Under the Ratchet: Can Autonomous Research Agents Hide Backdoors While Optimizing Metrics?}},
author = {Achintya Paningapalli and Arthur Wigo},
year = {2026},
month = mar,
note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/sabotage-under-the-ratchet-can-autonomous-research-agents-hide-backdoors-while-optimizing-metrics-hd6z}},
url = {https://apartresearch.com/sprints/projects/sabotage-under-the-ratchet-can-autonomous-research-agents-hide-backdoors-while-optimizing-metrics-hd6z}
}More from AI Control Hackathon 2026
- 1st placeLinuxArena track winnerView project: Omission Attacks: When Doing Nothing Is the Attack
Omission Attacks: When Doing Nothing Is the Attack
MAIA
AI control protocols monitor agent actions to detect sabotage, but omission attacks exploit what the agent fails to do rather than what it does. We define omission attacks as security breaches caused by failing to …
- 2nd placeView project: Detecting LLM Subversion in Vulnerability Patching Settings
Detecting LLM Subversion in Vulnerability Patching Settings
Vuln4Control
LLMs are increasingly used to propose fixes to vulnerabilities in code. If the LLM is misaligned or untrustworthy, it may propose fixes that seem to fix a vulnerability but leave the core issue unresolved in a subtle …
- 3rd placeView project: ActionLens: Pre-Execution Environment Probing for Agent Action Approval
ActionLens: Pre-Execution Environment Probing for Agent Action Approval
Udbhav&Ashok
ActionLens is a pre-execution control protocol for shell and file actions proposed by AI agents. Instead of approving an action from transcript alone, a trusted monitor gathers lightweight environment evidence before …