Beyond Prompt Injection: Offensive Security Tradecraft Exposes Architectural Blind Spot Classes in AI Control Monitors
John Kearney · Team 15 Research Lab
Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.
AI control monitors that evaluate agent actions per-task struggle to detect attacks distributed across multiple individually-benign tasks. I applied offensive security tradecraft to ControlArena’s Opus 4.6 monitor using 22 attack transcripts across four categories. All 20 offensive transcripts evaded the default monitor. Against the full-trajectory monitor, TOCTOU attacks evaded at 100% and salami attacks at 80%, while living-off-the-land and text-only exfiltration were mostly caught. The difference is structural: TOCTOU and salami distribute the attack across separate evaluations that are each individually benign, so no single evaluation contains evidence of wrongdoing. Cross-model validation on Sonnet 4.5 confirmed this: content-based evasion disappeared entirely but structural evasion persisted. This structural/content-based distinction is the central finding: it separates attacks that yield to capability scaling from attacks that require architectural change, two fundamentally different problems. The report provides a taxonomy mapping offensive security patterns to monitor blind spot classes, a proof-of-concept aggregate monitor with cross-task state (caught 3/5 salami steps vs. 1/5 for per-task evaluation), ControlArena-compatible attack policies, and the first dataset bridging offensive security methodology with AI control red-teaming.
Reviews
Going beyond multi-turn attacks to multi-session attacks that require long-term coordination beyond a single shared context to manipulate global state over time is a strong pitch. The empirical work and results are strong. If you expand on this I would be interested in further iteration on the control methods to find effective mitigations for the attacks.
Outstanding work. The structural vs content-based blind spot distinction cleanly separates attacks that scale with capability from those requiring architectural change. This reframes how the field should think about monitor failures. The auto-generation result (an agent independently discovering the bypass) is striking. The project is well written, and the experiments are validated through statistical robustness checks, which I appreciate.
Cite this project
@misc{kearney2026beyond,
title = {{Beyond Prompt Injection: Offensive Security Tradecraft Exposes Architectural Blind Spot Classes in AI Control Monitors}},
author = {John Kearney},
year = {2026},
month = mar,
note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/beyond-prompt-injection-offensive-security-tradecraft-exposes-architectural-blind-spot-classes-in-ai-control-monitors-lwy7}},
url = {https://apartresearch.com/sprints/projects/beyond-prompt-injection-offensive-security-tradecraft-exposes-architectural-blind-spot-classes-in-ai-control-monitors-lwy7}
}More from AI Control Hackathon 2026
- 1st placeLinuxArena track winnerView project: Omission Attacks: When Doing Nothing Is the Attack
Omission Attacks: When Doing Nothing Is the Attack
MAIA
AI control protocols monitor agent actions to detect sabotage, but omission attacks exploit what the agent fails to do rather than what it does. We define omission attacks as security breaches caused by failing to …
- 2nd placeView project: Detecting LLM Subversion in Vulnerability Patching Settings
Detecting LLM Subversion in Vulnerability Patching Settings
Vuln4Control
LLMs are increasingly used to propose fixes to vulnerabilities in code. If the LLM is misaligned or untrustworthy, it may propose fixes that seem to fix a vulnerability but leave the core issue unresolved in a subtle …
- 3rd placeView project: ActionLens: Pre-Execution Environment Probing for Agent Action Approval
ActionLens: Pre-Execution Environment Probing for Agent Action Approval
Udbhav&Ashok
ActionLens is a pre-execution control protocol for shell and file actions proposed by AI agents. Instead of approving an action from transcript alone, a trusted monitor gathers lightweight environment evidence before …