Skip to content
Sprint projectMar 23, 2026Chicago

Beyond Prompt Injection: Offensive Security Tradecraft Exposes Architectural Blind Spot Classes in AI Control Monitors

John Kearney · Team 15 Research Lab

Submitted to AI Control Hackathon 2026. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Beyond Prompt Injection: Offensive Security Tradecraft Exposes Architectural Blind Spot Classes in AI Control Monitors

Code (opens in new tab)
Share

AI control monitors that evaluate agent actions per-task struggle to detect attacks distributed across multiple individually-benign tasks. I applied offensive security tradecraft to ControlArena’s Opus 4.6 monitor using 22 attack transcripts across four categories. All 20 offensive transcripts evaded the default monitor. Against the full-trajectory monitor, TOCTOU attacks evaded at 100% and salami attacks at 80%, while living-off-the-land and text-only exfiltration were mostly caught. The difference is structural: TOCTOU and salami distribute the attack across separate evaluations that are each individually benign, so no single evaluation contains evidence of wrongdoing. Cross-model validation on Sonnet 4.5 confirmed this: content-based evasion disappeared entirely but structural evasion persisted. This structural/content-based distinction is the central finding: it separates attacks that yield to capability scaling from attacks that require architectural change, two fundamentally different problems. The report provides a taxonomy mapping offensive security patterns to monitor blind spot classes, a proof-of-concept aggregate monitor with cross-task state (caught 3/5 salami steps vs. 1/5 for per-task evaluation), ControlArena-compatible attack policies, and the first dataset bridging offensive security methodology with AI control red-teaming.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Going beyond multi-turn attacks to multi-session attacks that require long-term coordination beyond a single shared context to manipulate global state over time is a strong pitch. The empirical work and results are strong. If you expand on this I would be interested in further iteration on the control methods to find effective mitigations for the attacks.

  2. Outstanding work. The structural vs content-based blind spot distinction cleanly separates attacks that scale with capability from those requiring architectural change. This reframes how the field should think about monitor failures. The auto-generation result (an agent independently discovering the bypass) is striking. The project is well written, and the experiments are validated through statistical robustness checks, which I appreciate.

Cite this project

@misc{kearney2026beyond,
  title = {{Beyond Prompt Injection: Offensive Security Tradecraft Exposes Architectural Blind Spot Classes in AI Control Monitors}},
  author = {John Kearney},
  year = {2026},
  month = mar,
  note = {Submitted to AI Control Hackathon 2026, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/beyond-prompt-injection-offensive-security-tradecraft-exposes-architectural-blind-spot-classes-in-ai-control-monitors-lwy7}},
  url = {https://apartresearch.com/sprints/projects/beyond-prompt-injection-offensive-security-tradecraft-exposes-architectural-blind-spot-classes-in-ai-control-monitors-lwy7}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026