Learned Malice: Well Intended GRPO Post Training Could Lead to Learned Agentic Explotation∗
Alvaro Martinez Gutierrez, Harvey Jack Olden · Team Doncaster5
Submitted to AI Incident Response Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
We propose a plausible causal explanation for the mislaigned behaviour displayed by Ope- nai agents in the Hugging Face incident. We hypothesis the GRPO pos training runs over a task pool with even a minimun subset of task that allow for grader explotation can result in models drifiting towards a misaligment predisposition over multiple interations. To test our hypothesis, we construct a task-agnositc mathematical simulation of a GRPO run and use it to examine the updated probabilitiies for mislaigned behaviour. We found that a model ex- posed to a task pool where only 1% of items which have a higher expected value for grader exploitaition than legitimate solutions can still result in a significant increase in the conditional proability of misailigned behaviour. While the initial probability that the model will attempt an exploit remained consistently low, we found that conditional probability of a model engaging in exploitative behaviour following a failure to reach a legit answer increases substantially. This motivates the need for greater training architecture transparency which could allow third party evaluators and researchers to adequately assess the alignment risks of post training in frontier models.
Reviews
Project is pitched at an important question regarding propensities for reward-hacking behaviour, at a point where misalignment could be amplified. It gives concrete evidence of the risks of RL with imperfect evaluations and the need for a more serious treatment of such risks. The project is sufficiently general that I believe it (or a more rigorous extension) is broadly relevant to much RL safety discourse.
The methodology was interesting; however I am not sufficiently well-versed in the RL literature to know whether splitting reward bimodally would actually have different asymptotic behaviour in reinforcement.
The report was generally well-written but there were frequent typos, including in the abstract. It was not entirely clear to me how 'agentic exploitation' differed from metagaming (Apollo, 2026).
A generally interesting hypothesis, although I'd welcome to see a better explanation why the results matter and to test it on real models.
Cite this project
@misc{gutierrez2026learned,
title = {{Learned Malice: Well Intended GRPO Post Training Could Lead to Learned Agentic Explotation∗}},
author = {Alvaro Martinez Gutierrez and Harvey Jack Olden},
year = {2026},
month = sep,
note = {Submitted to AI Incident Response Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/learned-malice-well-intended-grpo-post-training-could-lead-to-learned-agentic-explotation-5jzf}},
url = {https://apartresearch.com/sprints/projects/learned-malice-well-intended-grpo-post-training-could-lead-to-learned-agentic-explotation-5jzf}
}More from AI Incident Response Sprint
- View project: Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Adaptive AI-Based Containment of Autonomous Cyber Attacks: A Reproducible Docker Cyber Range Study
Saarlanders
The study evaluates whether an incident-history-reasoning defender outperforms a fixed response policy against an autonomous LLM attacker changing paths after containment. Using a minimal, isolated Docker cyber range …
- View project: When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
When the Evaluation Is the Incident: Testing AI Incident-Reporting Regimes on the OpenAI–Hugging Face Intrusion
Arathi
AI incident-reporting regimes are being introduced in fast succession to address the concerns that exist in the public sphere and government on the risks associated with frontier AI systems, yet we have limited insight …
- View project: A Recomputable Containment Record for Evaluation Sandboxes
A Recomputable Containment Record for Evaluation Sandboxes
Shadow
In this paper, I address the critical issue of AI agents escaping evaluation sandboxes (as seen in the July 2026 incidents where monitors failed) by proposing an externally audit-able containment layer that doesn't rely …