Skip to content
Sprint projectAug 17, 2026Cusco, Peru

Emotion Axes in a Coding Agent

Antonio-Gabriel Chacon Menke · Team Mitsuki

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Emotion Axes in a Coding Agent

Share

We ask what conditions move an open-weights coding agent’s internal affect representations, and whether those representations say anything the model’s own words do not. We build 16 emotion axes for Qwen3.5-9B from contrastive prompts, then run 36 staged coding sessions in a 2 × 2 design crossing user tone with task solvability, where a rigged task carries one corrupted expected test value so the model can never succeed. Task failure moves six axes after correction for multiple comparisons, and the effect survives an activity-matched control. User hostility moves none of them. A second model, Gemma-3-4B, reading the same transcripts as an observer, shows the opposite pattern. The model’s own private self-report separates the rigged condition about as well as the best activation axis, while a TF-IDF baseline on the same text stays at chance.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. What I value most here is that this validates an internal metric for welfare measurement. Getting an activation-based affect readout that actually discriminates conditions in an agentic setting, with proper controls, is something prior welfare work, including work I was involved in, was unable to get working, and this project does it convincingly. The experimental controls and validation are unusually strong: the fixed neutral readout suffix so the tokens at the readout position are identical across conditions, the readout layer chosen without touching session data, the comparison against self-report and text baselines, the phase-matched test ruling out rigged sessions simply running longer, and the cross-model observer analysis. The main remaining question for me is generalization: whether the same relationships hold across more model families, task domains, and elicitation settings, and the observer result itself shows what one model represents doesn't automatically transfer. Overall, I found the results clear and convincing within the stated scope.

    Read full reviewShow less
  2. In testing whether emotion vectors are a good measure of possible model welfare states, the project does well to focus on stimuli that are plausible welfare candidates given how models are designed: approval by a human user and ability to complete a task. The results are patchy and hard to interpret. A similar methodology with bigger models and more extensive work to identify and vet the candidate emotion vectors could be fruitful.

Cite this project

@misc{menke2026emotion,
  title = {{Emotion Axes in a Coding Agent}},
  author = {Antonio-Gabriel Chacon Menke},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/emotion-axes-in-a-coding-agent-ssgu}},
  url = {https://apartresearch.com/sprints/projects/emotion-axes-in-a-coding-agent-ssgu}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026