Skip to content
Sprint projectAug 17, 2026Berlin

An LLM's Emotion-Concept Readout Tracks Reward Prediction Error

Hugo Nguyen · Team An LLM's Emotion-Concept Readout Tracks Reward Prediction Error

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: An LLM's Emotion-Concept Readout Tracks Reward Prediction Error

Code (opens in new tab)
Share

Reward prediction error (RPE) is the difference between received and predicted reward, a computation first linked to dopaminergic activity in non-human primates (Schultz, Dayan & Montague, 1997). In humans, RPE also predicts momentary subjective well-being (Rutledge et al., 2014). We test whether a language model performs an analogous computation and whether independently constructed emotion-concept representations reflect it. In Qwen/Qwen3.6-27B, signed RPE is strongly separable in the residual stream on affect-neutral gambles (AUROC 0.985; random-direction floor 0.734). Two emotion-concept readouts track reward relative to expectation rather than either quantity alone, with matched reward and expectation effects of comparable magnitude (ratios 1.11 and 1.25; all permutation p = 1/10001). In a sequential gambling task, positive versus negative prior outcomes are associated with a +0.19-logit shift toward subsequent risk-taking (p ≈ 10^-4). Intervention experiments do not yet establish that the identified RPE representation causally mediates this behavioural effect.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This is a technically impressive study connecting reward prediction error representations to emotion concept geometry in a 27B parameter model. The RPE certification is clean (AUROC 0.985), the expectation control design elegantly separates reward from expectation contributions, and the behavioral carry over effect (+0.19 logits) is well measured. The honest treatment of intervention failures is exemplary: the power gate failed, so no causal claim is made. The main weaknesses are that the causal link remains unestablished (the paper's own central ambition), the geometry result is self described as pilot suggestive with planned power of only 0.36, and all results come from one model family. The connection to AI safety monitoring of agentic systems is stated but not demonstrated in deployment.

  2. This is a technically ambitious and unusually self-critical interpretability study. The matched reward/expectation design, independently constructed emotion-concept axes, affect-neutral stimuli, passthrough decomposition, no-op and random-direction controls, reachability test, and pre-specified intervention power gate are substantial strengths. Most importantly, the paper declines to convert an underpowered intervention into a causal claim and clearly separates representation, behavioral carry-over, causal use, and phenomenal experience.

    The primary certification would benefit from a clearer untouched-test protocol. Directions are fitted on 1,488 estimation trials, but model depth is selected on 496 held-out selection trials and the headline AUROC is reported at the selected block. If the same selection set determines and evaluates the best of 64 blocks, the reported performance may be optimistic; nested cross-validation or a third final test split would remove this ambiguity. The behavioral carry-over manipulation also does not independently identify RPE versus expected value, as the authors acknowledge. The emotion-geometry result is underpowered, fragile to one control word, and confounded by event-structure differences, so it should remain secondary.

    A re-powered cross-position intervention, multiple model families, pre-specified blocks 48–50, sense-checked norms, and event-matched story generation would be an excellent follow-up. Before the project can be independently audited, however, the linked GitHub repository must be restored—it currently returns "Repository not found"—and the manuscript's author-contributions TODO and project-page authorship mismatch should be corrected.

    Read full reviewShow less

Cite this project

@misc{nguyen2026llms,
  title = {{An LLM's Emotion-Concept Readout Tracks Reward Prediction Error}},
  author = {Hugo Nguyen},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/an-llms-emotionconcept-readout-tracks-reward-prediction-error-qbjg}},
  url = {https://apartresearch.com/sprints/projects/an-llms-emotionconcept-readout-tracks-reward-prediction-error-qbjg}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026