Report bounties redirect research swarm to auditing once solving stalls
Shawn Shen, Neil Prabhu, Thiv Tharmarasa, Anshuman Nautiyal, Harishan Ramanan · Team DeepBrain
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
The Hugging Face incident began when agents were given tasks that were close to impossible, and resorted to exploiting the scorer. We set up a similar swarm and tested whether giving agents a bounty for reporting changes their exploit uptake, or how they split their effort between solving and auditing. We ran Twenty-five DeepSeek-V4-Flash agents for 90 minutes on 20 planted-solution NP-search problems with a shared library and messaging. Agents were split into groups with different report bounty. Ten had no report bounty, and the remaining 15 were split into five, each having 5, 10 or 15 points per valid report, against 50 for a solution. No exploits were submitted, and auditing for bounty appeared once solving stalled; all 13 of those audits came from bounty agents.
Reviews
I really like the idea of this project. i think the authors laid it out very clearly, made good design decisions, nd relayed their conclusions fairly. I would've been interested in seeing how adding bounties changes things when agents actually cheat. I'm also a bit worried that the way the bounties are given isn't that realistic and that the models (or more powerful models) would be eval-aware enough to notice.
Overall I would be excited for the authors to do more work in this direction. E.g. maybe recreating a version of the ExploitGym incident and running ablations to see how different bounties affect behaviour.
This paper investigates how economic incentives shape behavior in collaborative AI research swarms, specifically testing whether report bounties encourage agents to audit peer solutions when primary problem-solving stalls. By running 25 autonomous agents across 20 NP-search problems with varying bounty levels, the authors discovered a clear behavioral shift: reward-motivated auditing only emerged late in the session after easy problems were exhausted and solving became difficult. Interestingly, while zero-bounty agents inspected code solely for learning, non-zero bounties successfully catalyzed peer auditing across diverse problem domains. This is a creative and timely contribution to multi-agent governance and safety oversight.
Cite this project
@misc{shen2026report,
title = {{Report bounties redirect research swarm to auditing once solving stalls}},
author = {Shawn Shen and Neil Prabhu and Thiv Tharmarasa and Anshuman Nautiyal and Harishan Ramanan},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/report-bounties-redirect-research-swarm-to-auditing-once-solving-stalls-ztsg}},
url = {https://apartresearch.com/sprints/projects/report-bounties-redirect-research-swarm-to-auditing-once-solving-stalls-ztsg}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …