Emotion Axes in a Coding Agent
Antonio-Gabriel Chacon Menke
We ask what conditions move an open-weights coding agent’s internal affect representations, and whether those representations say anything the model’s own words do not. We build 16 emotion axes for Qwen3.5-9B from contrastive prompts, then run 36 staged coding sessions in a 2 × 2 design crossing user tone with task solvability, where a rigged task carries one corrupted expected test value so the model can never succeed. Task failure moves six axes after correction for multiple comparisons, and the effect survives an activity-matched control. User hostility moves none of them. A second model, Gemma-3-4B, reading the same transcripts as an observer, shows the opposite pattern. The model’s own private self-report separates the rigged condition about as well as the best activation axis, while a TF-IDF baseline on the same text stays at chance.
What I value most here is that this validates an internal metric for welfare measurement. Getting an activation-based affect readout that actually discriminates conditions in an agentic setting, with proper controls, is something prior welfare work, including work I was involved in, was unable to get working, and this project does it convincingly. The experimental controls and validation are unusually strong: the fixed neutral readout suffix so the tokens at the readout position are identical across conditions, the readout layer chosen without touching session data, the comparison against self-report and text baselines, the phase-matched test ruling out rigged sessions simply running longer, and the cross-model observer analysis. The main remaining question for me is generalization: whether the same relationships hold across more model families, task domains, and elicitation settings, and the observer result itself shows what one model represents doesn't automatically transfer. Overall, I found the results clear and convincing within the stated scope.
In testing whether emotion vectors are a good measure of possible model welfare states, the project does well to focus on stimuli that are plausible welfare candidates given how models are designed: approval by a human user and ability to complete a task. The results are patchy and hard to interpret. A similar methodology with bigger models and more extensive work to identify and vet the candidate emotion vectors could be fruitful.
Cite this work
@misc {
title={
(HckPrj) Emotion Axes in a Coding Agent
},
author={
Antonio-Gabriel Chacon Menke
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


