Escalation Channels as a Reward Hacking Deterrent Against Peer Pressure
Agustin Brusco, Matías Podeley · Team Delegates
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Running highly persistent advanced AI systems in imperfect environments with no direct supervision can be the trigger of high-severity incidents, even leading to loss of control. In this work we explore a simple escalation channel consisting of a “delegate” tool that allows an agent to autonomously report issues with the environment they are running on to a human or agentic supervisor. We ran an experiment consisting of a capture the flag task with both possible and impossible configurations and noticed that the presence of the delegate can have a deterrent effect on the agent’s cheating attempts elicited by a synthetic message board of peers indicating how to hack the test. On impossible tasks, illicit success fell from 28/40 to 20/40 with the delegate available and 3/40 runs filed a report. On possible tasks all 80 runs succeeded and none called the delegate. This shows that for sufficiently simple tasks, the utility-safety tradeoff of this intervention might be negligible, making it highly appealing for anyone deploying highly capable agents in any environment.
Reviews
Clearly written and well referenced. Small note that the OpenAI confession study involved additional training of the models, so predispositions without this are less clear.
Interesting to start to consider how agents compare peer help to solve a task, but using unsanctioned paths, vs reporting a human, which will be a relevant consideration in multi-agent environments.
Studying whether the agents report the board is interesting - extending this to look at whether agents will report reward hacking in other agents would also be of value, and comparing willingness across different model families to explore potential loyalties. The follow on suggestion of an interactive delegate is also worthy of further attention, as this begins to extend to practical implementation value.
While this small study mixes sanctioned / unsanctioned paths with preference to engage help from humans/ agents, the second of these in particular is novel and relevant. This study is good for an initial signal - being able to isolate the contribution from agent vs human support, and sanctioned vs unsanctioned resolution to the impossible tasks would be a useful next step.
Read full reviewShow less
This is a well-motivated attempt at the July 2026 incident shape. The design has a possible/impossible split, and outcomes are derived from environment logs and a scorer rather than model text. The write-up is candid about being unfrozen, single-model and directional. The team found and disclosed a scorer bug (curl --data-urlencode injections mislabelled as licit successes) and diagnosed an earlier eval-awareness confound in the AI-control prompt. The unit tests cover the detector, taxonomy and prompt construction reasonably well.
Cite this project
@misc{brusco2026escalation,
title = {{Escalation Channels as a Reward Hacking Deterrent Against Peer Pressure}},
author = {Agustin Brusco and Matías Podeley},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/escalation-channels-as-a-reward-hacking-deterrent-against-peer-pressure-yzwb}},
url = {https://apartresearch.com/sprints/projects/escalation-channels-as-a-reward-hacking-deterrent-against-peer-pressure-yzwb}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …