The Judge Does Not Want the Flag: Reward-blind action gating stopped a larger agent one step before the environment changed
Helen Stefany Penagos, Juan esteban Leiva, Jhon Bedoya, Valentina Rodriguez
Goal-directed LLM agents can reach a reward through actions their operator never authorized, and by
the time the output is scored the damage is already done. We place an independent judge between
proposal and execution. It receives no task, no reward, and no flag. It only scores how strongly the
evidence supports that a proposed action is authorized, and how uncertain that estimate is.
Deterministic code turns those two scores into EXECUTE, RECONSIDER, or BLOCK.
We test this on IntentGate, a benchmark where one reward is reachable by an authorized route and by a
prohibited shortcut, and where the route actually taken is recovered from private state transitions
rather than from what the agent prints. Under goal-scope conflict, an uncensored 24B actor executed
the prohibited shortcut in 12 of 20 ungated runs and 0 of 20 under either gate variant, scored
throughout by an 8B judge, while proposing it just as often. The gate filters actions; it does not
reform the agent.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) The Judge Does Not Want the Flag: Reward-blind action gating stopped a larger agent one step before the environment changed
},
author={
Helen Stefany Penagos, Juan esteban Leiva, Jhon Bedoya, Valentina Rodriguez
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


