An LLM's Emotion-Concept Readout Tracks Reward Prediction Error
Hugo Nguyen
Reward prediction error (RPE) is the difference between received and predicted reward, a computation first linked to dopaminergic activity in non-human primates (Schultz, Dayan & Montague, 1997). In humans, RPE also predicts momentary subjective well-being (Rutledge et al., 2014). We test whether a language model performs an analogous computation and whether independently constructed emotion-concept representations reflect it. In Qwen/Qwen3.6-27B, signed RPE is strongly separable in the residual stream on affect-neutral gambles (AUROC 0.985; random-direction floor 0.734). Two emotion-concept readouts track reward relative to expectation rather than either quantity alone, with matched reward and expectation effects of comparable magnitude (ratios 1.11 and 1.25; all permutation p = 1/10001). In a sequential gambling task, positive versus negative prior outcomes are associated with a +0.19-logit shift toward subsequent risk-taking (p ≈ 10^-4). Intervention experiments do not yet establish that the identified RPE representation causally mediates this behavioural effect.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) An LLM's Emotion-Concept Readout Tracks Reward Prediction Error
},
author={
Hugo Nguyen
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


