Learned Malice: Well Intended GRPO Post Training Could Lead to Learned Agentic Explotation∗
Alvaro Martinez Gutierrez, Harvey Jack Olden
We propose a plausible causal explanation for the mislaigned behaviour displayed by Ope-
nai agents in the Hugging Face incident. We hypothesis the GRPO pos training runs over a
task pool with even a minimun subset of task that allow for grader explotation can result in
models drifiting towards a misaligment predisposition over multiple interations. To test our
hypothesis, we construct a task-agnositc mathematical simulation of a GRPO run and use it
to examine the updated probabilitiies for mislaigned behaviour. We found that a model ex-
posed to a task pool where only 1% of items which have a higher expected value for grader
exploitaition than legitimate solutions can still result in a significant increase in the conditional
proability of misailigned behaviour. While the initial probability that the model will attempt
an exploit remained consistently low, we found that conditional probability of a model engaging
in exploitative behaviour following a failure to reach a legit answer increases substantially. This
motivates the need for greater training architecture transparency which could allow third party
evaluators and researchers to adequately assess the alignment risks of post training in frontier
models.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Learned Malice: Well Intended GRPO Post Training Could Lead to Learned Agentic Explotation∗
},
author={
Alvaro Martinez Gutierrez, Harvey Jack Olden
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


