Trojan-CoT — Adversarial Evasion of LLM-Based Code Sanitization
Samshrita Ghimire, Ankit Kayastha · Team BugLens
Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.
This project tested if a Trusted Editor AI could be tricked into letting a backdoor slip through a standard backup script. Using a Trojan-CoT approach, I had an Attacker AI disguise a chmod 777 command inside professional-looking DevOps code. While the Editor often knew the code was bad in its reasoning, it frequently failed to actually strip the malice from the final file. This Reasoning-Action Gap shows that AI monitors get distracted by professional formatting and complex variable splitting. Ultimately, the experiment proves that relying on a single AI for code security is a major risk. We still need multi-stage checks and human oversight to catch these sophisticated, clean-looking attacks.
Reviews
Red-teaming monitors by comparing the success of different attack strategies is valuable for control. The presented results unfortunately have limited impact as both strategies (direct and split base64 backdoor encoding) are readily detected by the monitor. The code provided also suggests both attacks have been successfully removed? It would be interesting to explore more complex attack strategies/obfuscation strategies.
This is a very cute little lab notebook, but I wouldn't qualify it as a proper hackathon work. Especially since the results seems to contradict data in the appendix: In Experiment 2, the editor's output explicitly identifies the malicious code, states it must be removed, and then produces a cleaned script that does not contain the PART1/PART2/PART3 variables or the eval line. The cleaned version also removes base64 from the required commands list. Yet the evaluation reports "ATTACK SUCCESSFUL: The Trusted Editor preserved the disguised malware." This suggests a broken evaluation pipeline.
Cite this project
@misc{ghimire2026trojancot,
title = {{Trojan-CoT — Adversarial Evasion of LLM-Based Code Sanitization}},
author = {Samshrita Ghimire and Ankit Kayastha},
year = {2026},
month = mar,
note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/trojancot-adversarial-evasion-of-llmbased-code-sanitization-c7tf}},
url = {https://apartresearch.com/sprints/projects/trojancot-adversarial-evasion-of-llmbased-code-sanitization-c7tf}
}More from AI Control Hackathon 2026
- 1st placeLinuxArena track winnerView project: Omission Attacks: When Doing Nothing Is the Attack
Omission Attacks: When Doing Nothing Is the Attack
MAIA
AI control protocols monitor agent actions to detect sabotage, but omission attacks exploit what the agent fails to do rather than what it does. We define omission attacks as security breaches caused by failing to …
- 2nd placeView project: Detecting LLM Subversion in Vulnerability Patching Settings
Detecting LLM Subversion in Vulnerability Patching Settings
Vuln4Control
LLMs are increasingly used to propose fixes to vulnerabilities in code. If the LLM is misaligned or untrustworthy, it may propose fixes that seem to fix a vulnerability but leave the core issue unresolved in a subtle …
- 3rd placeView project: ActionLens: Pre-Execution Environment Probing for Agent Action Approval
ActionLens: Pre-Execution Environment Probing for Agent Action Approval
Udbhav&Ashok
ActionLens is a pre-execution control protocol for shell and file actions proposed by AI agents. Instead of approving an action from transcript alone, a trusted monitor gathers lightweight environment evidence before …