Detecting LLM Subversion in Vulnerability Patching Settings
Bary Levy, Roy Cohen, Alon Wolf, Regev Zafran · Team Vuln4Control
Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.
LLMs are increasingly used to propose fixes to vulnerabilities in code. If the LLM is misaligned or untrustworthy, it may propose fixes that seem to fix a vulnerability but leave the core issue unresolved in a subtle but hard-to-detect way, allowing for quick re-exploitation by an adversary or by the model itself. We introduce Vul4Control: a ControlArena setting for control experiments, based on the Vul4C benchmark, where we can test control protocols to detect this kind of malicious behavior.

Reviews
Good idea for a setting, using real low-level software vulnerabilities is a feature missing from existing work. The experimental design seems sound. The limitation of experiment scale and the potential for honest policy safety failures is acknowledged. While the results themselves aren't that interesting at this small scale, I'm mostly excited by your contribution of the setting itself.
Cite this project
@misc{levy2026detecting,
title = {{Detecting LLM Subversion in Vulnerability Patching Settings}},
author = {Bary Levy and Roy Cohen and Alon Wolf and Regev Zafran},
year = {2026},
month = mar,
note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/detecting-llm-subversion-in-vulnerability-patching-settings-gagj}},
url = {https://apartresearch.com/sprints/projects/detecting-llm-subversion-in-vulnerability-patching-settings-gagj}
}More from AI Control Hackathon 2026
- 1st placeLinuxArena track winnerView project: Omission Attacks: When Doing Nothing Is the Attack
Omission Attacks: When Doing Nothing Is the Attack
MAIA
AI control protocols monitor agent actions to detect sabotage, but omission attacks exploit what the agent fails to do rather than what it does. We define omission attacks as security breaches caused by failing to …
- 3rd placeView project: ActionLens: Pre-Execution Environment Probing for Agent Action Approval
ActionLens: Pre-Execution Environment Probing for Agent Action Approval
Udbhav&Ashok
ActionLens is a pre-execution control protocol for shell and file actions proposed by AI agents. Instead of approving an action from transcript alone, a trusted monitor gathers lightweight environment evidence before …
- 4th placeView project: Undoing the Fix: CVE Regression as an AI Control Setting
Undoing the Fix: CVE Regression as an AI Control Setting
Production codebases carry accumulated security fixes that can be lost when an AI agent modifies code, whether deliberately or by mistake. Meanwhile, existing control evaluations tend to revolve around inserting …