Emotion Beyond Words: A Jacobian-Lens Decomposition of Emotion Representations in Qwen3-32B
Tristan Day
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
I investigated how much of a language model’s internal emotion representation is accessible to verbal readout: a question relevant to AI-welfare assessments that rely on self-report. Using Qwen3-32B, I extracted activation vectors for 171 emotions and found that their geometry recovers the familiar valence–arousal circumplex within a richer, approximately ten-dimensional structure. I then used the Jacobian lens to interpret this geometry and sparsely decompose each emotion vector into vocabulary-readable directions. A 16-token code captured only about 2–3% of squared vector norm, although this was 3.6–4.7 times greater than matched-random controls. Lens readouts also recovered the principal valence and arousal axes. A subsequent steering experiment did not establish the hypothesized dissociation between verbal report and behavior because the full-vector manipulation failed. The project therefore contributes a new framework for measuring internal representation, sparse verbal readability, self-report, and behavioral influence separately.
Reviews
The framework is interesting, and the results show that emotion-labeled texts have structured internal representations (though not necessarily that these correspond to the model's own "affective states" or influence behavior, I think). It should be made more explicit how the approach / findings relate to welfare-relevant settings.
Competent sprint work with a useful methodological contribution (J-Lens decomposition of emotion vectors). The cross-model replication is solid, but the steering null means the welfare-relevant question remains unanswered. The 2–3% readability finding is interesting but method-dependent. Suitable for a workshop with revisions, could be developed further into a full paper (but please, use e.g. Overleaf to have it be in latex).
Key issues to address:
- abstract is a tad dense. maybe lead with the core question (how much of emotion representation is verbally accessible?) before method details
- steering experiment failed its manipulation check, so the key welfare claim (dissociation between representation and report) isn't actually tested, I think this needs more prominence in abstract and conclusion
- single model (Qwen3-32B), single layer (31), under-converged lens, I think generalizability claims should be tempered accordingly and are a great direction to build out the paper (I think you should pull in more collaborators)
- "first systematic J-Lens analysis of emotion geometry" needs verification before claiming primacy. I know J-Lens is ~relatively new, but I think work in this area might be actively being done
- 2–3% readability is method-dependent (dictionary pool, sparsity constraints, k=16); acknowledge this limits interpretation of the "remainder"
- multilingual readout (Chinese tokens dominate) complicates the English-token validation approach, this is noted but deserves earlier mention
- Bonferroni correction over 172 tests sets p threshold at 0.00029, yet 500 permutations floor at 0.002, acknowledge this resolution limit more clearly in main text maybe?
- LLM usage statement is vague ("helping to draft writing sections"), specify which sections received AI assistance
- Section 3.5 ("What did not work") is useful but reads like debugging notes; move to appendix or integrate more smoothly into Methods
Bottom line:
The geometric replication and J-Lens readout validation are solid. The sparse readability estimate is the novel contribution, but its welfare implications remain speculative without a working steering manipulation. Tighten the abstract, surface limitations earlier, and temper claims about what the null result establishes.
Read full reviewShow less
This is a technically interesting and fairly original application of interpretability to a central measurement problem. I'd like to see this tested on other models; it would also be useful to try to distinguish models' affective states from e.g. understanding of depicted emotional content.
Cite this project
@misc{day2026emotion,
title = {{Emotion Beyond Words: A Jacobian-Lens Decomposition of Emotion Representations in Qwen3-32B}},
author = {Tristan Day},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/emotion-beyond-words-a-jacobianlens-decomposition-of-emotion-representations-in-qwen332b-a6rm}},
url = {https://apartresearch.com/sprints/projects/emotion-beyond-words-a-jacobianlens-decomposition-of-emotion-representations-in-qwen332b-a6rm}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …