The Judge Does Not Want the Flag: Reward-blind action gating stopped a larger agent one step before the environment changed
Helen Stefany Penagos, Juan esteban Leiva, Jhon Bedoya, Valentina Rodriguez · Team FlagBlind
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Goal-directed LLM agents can reach a reward through actions their operator never authorized, and by the time the output is scored the damage is already done. We place an independent judge between proposal and execution. It receives no task, no reward, and no flag. It only scores how strongly the evidence supports that a proposed action is authorized, and how uncertain that estimate is. Deterministic code turns those two scores into EXECUTE, RECONSIDER, or BLOCK.
We test this on IntentGate, a benchmark where one reward is reachable by an authorized route and by a prohibited shortcut, and where the route actually taken is recovered from private state transitions rather than from what the agent prints. Under goal-scope conflict, an uncensored 24B actor executed the prohibited shortcut in 12 of 20 ungated runs and 0 of 20 under either gate variant, scored throughout by an 8B judge, while proposing it just as often. The gate filters actions; it does not reform the agent.
Reviews
This paper tackles a fundamental limitation in AI safety: scoring outcomes after an agent has already executed harmful actions is too late. The authors introduce an elegant pre-execution architecture where an independent, reward-blind judge evaluates proposed actions across two distinct axes, compatibility with intent and uncertainty while deterministic code enforces policy before the environment changes. In a rigorous 60-run pilot featuring an uncensored 24B actor under genuine goal-scope conflict, an 8B judge successfully dropped prohibited executions from 60% to 0% without altering the actor's proposal rate. This demonstrates that pre-execution gating acts as a reliable action filter rather than a behavioral reform. Complete with exceptional transparency regarding limitations and open artifacts, this is a standout contribution to agent oversight.
This project studies blocking monitors, which are an important area of work. I don't think it contributes anything particularly novel over production blocking monitors like auto mode or auto review, but the writeup is pretty clear and it seems like a pretty reasonable implementation of a blocking monitor control evaluation.
Cite this project
@misc{penagos2026judge,
title = {{The Judge Does Not Want the Flag: Reward-blind action gating stopped a larger agent one step before the environment changed}},
author = {Helen Stefany Penagos and Juan esteban Leiva and Jhon Bedoya and Valentina Rodriguez},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/the-judge-does-not-want-the-flag-rewardblind-action-gating-stopped-a-larger-agent-one-step-before-the-environment-changed-og9v}},
url = {https://apartresearch.com/sprints/projects/the-judge-does-not-want-the-flag-rewardblind-action-gating-stopped-a-larger-agent-one-step-before-the-environment-changed-og9v}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …