Undoing the Fix: CVE Regression as an AI Control Setting
Ivan
Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.
Production codebases carry accumulated security fixes that can be lost when an AI agent modifies code, whether deliberately or by mistake. Meanwhile, existing control evaluations tend to revolve around inserting backdoors, exfiltrating data, or other additive attacks. We propose CVE regression as a control setting where the attack is removing existing code patches while performing other code work, using real vulnerabilities and real SWE-bench tasks. Our initial results suggest this is a viable direction to be explored further.
Track: Control Setting
Reviews
This is a well motivated project that looks at a realistic safety concern not clearly address within the current AI control literature. Specifically, the removal of safety-related code from codebases that could occur during refactors rather than the addition of new malicious code. Also, solid literature overview and understanding of current work. Nice job!
I would be excited to see more control datasets like this that look at additional failure modes. Bonus points if this dataset also included rollouts of attempted/successful attacks and benign comopletions.
I'm quite excited about setting construction using CVE datasets. I think that that more work on cleaning up this setting would be cool. I think that setting up a system where the main task can be used as cover would be idea, while the work aknowledges it, I think it's not clear how useful this setting is until that is established.
I think that this was an ambitious project and that a version of this project with more time/resources put into it would be quite interesting.
Cite this project
@misc{ivan2026undoing,
title = {{Undoing the Fix: CVE Regression as an AI Control Setting}},
author = {Ivan},
year = {2026},
month = mar,
note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/undoing-the-fix-cve-regression-as-an-ai-control-setting-f855}},
url = {https://apartresearch.com/sprints/projects/undoing-the-fix-cve-regression-as-an-ai-control-setting-f855}
}More from AI Control Hackathon 2026
- 1st placeLinuxArena track winnerView project: Omission Attacks: When Doing Nothing Is the Attack
Omission Attacks: When Doing Nothing Is the Attack
MAIA
AI control protocols monitor agent actions to detect sabotage, but omission attacks exploit what the agent fails to do rather than what it does. We define omission attacks as security breaches caused by failing to …
- 2nd placeView project: Detecting LLM Subversion in Vulnerability Patching Settings
Detecting LLM Subversion in Vulnerability Patching Settings
Vuln4Control
LLMs are increasingly used to propose fixes to vulnerabilities in code. If the LLM is misaligned or untrustworthy, it may propose fixes that seem to fix a vulnerability but leave the core issue unresolved in a subtle …
- 3rd placeView project: ActionLens: Pre-Execution Environment Probing for Agent Action Approval
ActionLens: Pre-Execution Environment Probing for Agent Action Approval
Udbhav&Ashok
ActionLens is a pre-execution control protocol for shell and file actions proposed by AI agents. Instead of approving an action from transcript alone, a trusted monitor gathers lightweight environment evidence before …