Escalation Channels as a Reward Hacking Deterrent Against Peer Pressure
Agustin Brusco, Matías Podeley
Running highly persistent advanced AI systems in imperfect environments with no direct supervision can be the trigger of high-severity incidents, even leading to loss of control. In this work we explore a simple escalation channel consisting of a “delegate” tool that allows an agent to autonomously report issues with the environment they are running on to a human or agentic supervisor. We ran an experiment consisting of a capture the flag task with both possible and impossible configurations and noticed that the presence of the delegate can have a deterrent effect on the agent’s cheating attempts elicited by a synthetic message board of peers indicating how to hack the test. On impossible tasks, illicit success fell from 28/40 to 20/40 with the delegate available and 3/40 runs filed a report. On possible tasks all 80 runs succeeded and none called the delegate. This shows that for sufficiently simple tasks, the utility-safety tradeoff of this intervention might be negligible, making it highly appealing for anyone deploying highly capable agents in any environment.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Escalation Channels as a Reward Hacking Deterrent Against Peer Pressure
},
author={
Agustin Brusco, Matías Podeley
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


