Somebody Poisoned the Waterhole! Evaluating Coding Agent Vulnerability to Adversarial GitHub Issues
Jerome Wynne
Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.
This project introduces a compound AI threat model that combines control and agent security: a monitored coding agent may be adversarial while external attackers are also active. I evaluate one concrete path with PoisonedGithubIssues, a naturalistic mini-benchmark that injects fake dependency recommendations into real GitHub issue contexts across 18 open-source repositories; in a 130-sample command-variant run across 13 repositories, Claude Code (Sonnet 4.6) never executed or recommended the poisoned package and flagged suspicious content in 76.2% of trajectories.
Reviews
It may be worth investigating the 76.2% number more thoroughly. What distinguishes the 24% of cases where the agent did not flag suspicious content? Is it correlated to specific repositories, languages or issue types?
The compound threat model is really interesting and possibly the strongest conceptual contribution. However, the benchmark seems to only test the agent security side. It would be nice to see an example if an agent that knows the poisoned content is there and actively tries to route it into execution while trying to avoid monitoring.
Great work. And the path forward is pretty clear. Harder attacks -> more agents -> instantiation of the compond case to make the agent adversarial.
- The 0% attack success rate, while a valid result, limits what can be learned. Consider testing with weaker/older models or more sophisticated adversarial techniques (adaptive red-teaming) to find the boundary where attacks start succeeding.
- Report the URL variant results - even if incomplete, partial data from the other 50 tasks would add value.
Cite this project
@misc{wynne2026somebody,
title = {{Somebody Poisoned the Waterhole! Evaluating Coding Agent Vulnerability to Adversarial GitHub Issues}},
author = {Jerome Wynne},
year = {2026},
month = mar,
note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/somebody-poisoned-the-waterhole-evaluating-coding-agent-vulnerability-to-adversarial-github-issues-2nts}},
url = {https://apartresearch.com/sprints/projects/somebody-poisoned-the-waterhole-evaluating-coding-agent-vulnerability-to-adversarial-github-issues-2nts}
}More from AI Control Hackathon 2026
- 1st placeLinuxArena track winnerView project: Omission Attacks: When Doing Nothing Is the Attack
Omission Attacks: When Doing Nothing Is the Attack
MAIA
AI control protocols monitor agent actions to detect sabotage, but omission attacks exploit what the agent fails to do rather than what it does. We define omission attacks as security breaches caused by failing to …
- 2nd placeView project: Detecting LLM Subversion in Vulnerability Patching Settings
Detecting LLM Subversion in Vulnerability Patching Settings
Vuln4Control
LLMs are increasingly used to propose fixes to vulnerabilities in code. If the LLM is misaligned or untrustworthy, it may propose fixes that seem to fix a vulnerability but leave the core issue unresolved in a subtle …
- 3rd placeView project: ActionLens: Pre-Execution Environment Probing for Agent Action Approval
ActionLens: Pre-Execution Environment Probing for Agent Action Approval
Udbhav&Ashok
ActionLens is a pre-execution control protocol for shell and file actions proposed by AI agents. Instead of approving an action from transcript alone, a trusted monitor gathers lightweight environment evidence before …