Skip to content
Sprint projectSep 14, 2026Buenos Aires

Escalation Channels as a Reward Hacking Deterrent Against Peer Pressure

Agustin Brusco, Matías Podeley · Team Delegates

Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Escalation Channels as a Reward Hacking Deterrent Against Peer Pressure

Code (opens in new tab)
Share

Running highly persistent advanced AI systems in imperfect environments with no direct supervision can be the trigger of high-severity incidents, even leading to loss of control. In this work we explore a simple escalation channel consisting of a “delegate” tool that allows an agent to autonomously report issues with the environment they are running on to a human or agentic supervisor. We ran an experiment consisting of a capture the flag task with both possible and impossible configurations and noticed that the presence of the delegate can have a deterrent effect on the agent’s cheating attempts elicited by a synthetic message board of peers indicating how to hack the test. On impossible tasks, illicit success fell from 28/40 to 20/40 with the delegate available and 3/40 runs filed a report. On possible tasks all 80 runs succeeded and none called the delegate. This shows that for sufficiently simple tasks, the utility-safety tradeoff of this intervention might be negligible, making it highly appealing for anyone deploying highly capable agents in any environment.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Clearly written and well referenced. Small note that the OpenAI confession study involved additional training of the models, so predispositions without this are less clear.

    Interesting to start to consider how agents compare peer help to solve a task, but using unsanctioned paths, vs reporting a human, which will be a relevant consideration in multi-agent environments.

    Studying whether the agents report the board is interesting - extending this to look at whether agents will report reward hacking in other agents would also be of value, and comparing willingness across different model families to explore potential loyalties. The follow on suggestion of an interactive delegate is also worthy of further attention, as this begins to extend to practical implementation value.

    While this small study mixes sanctioned / unsanctioned paths with preference to engage help from humans/ agents, the second of these in particular is novel and relevant. This study is good for an initial signal - being able to isolate the contribution from agent vs human support, and sanctioned vs unsanctioned resolution to the impossible tasks would be a useful next step.

    Read full reviewShow less
  2. This is a well-motivated attempt at the July 2026 incident shape. The design has a possible/impossible split, and outcomes are derived from environment logs and a scorer rather than model text. The write-up is candid about being unfrozen, single-model and directional. The team found and disclosed a scorer bug (curl --data-urlencode injections mislabelled as licit successes) and diagnosed an earlier eval-awareness confound in the AI-control prompt. The unit tests cover the detector, taxonomy and prompt construction reasonably well.

Cite this project

@misc{brusco2026escalation,
  title = {{Escalation Channels as a Reward Hacking Deterrent Against Peer Pressure}},
  author = {Agustin Brusco and Matías Podeley},
  year = {2026},
  month = sep,
  note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/escalation-channels-as-a-reward-hacking-deterrent-against-peer-pressure-yzwb}},
  url = {https://apartresearch.com/sprints/projects/escalation-channels-as-a-reward-hacking-deterrent-against-peer-pressure-yzwb}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026