Skip to content
Sprint projectAug 17, 2026Cambridge, MA

Emotion Beyond Words: A Jacobian-Lens Decomposition of Emotion Representations in Qwen3-32B

Tristan Day

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Emotion Beyond Words: A Jacobian-Lens Decomposition of Emotion Representations in Qwen3-32B

Code (opens in new tab)
Share

I investigated how much of a language model’s internal emotion representation is accessible to verbal readout: a question relevant to AI-welfare assessments that rely on self-report. Using Qwen3-32B, I extracted activation vectors for 171 emotions and found that their geometry recovers the familiar valence–arousal circumplex within a richer, approximately ten-dimensional structure. I then used the Jacobian lens to interpret this geometry and sparsely decompose each emotion vector into vocabulary-readable directions. A 16-token code captured only about 2–3% of squared vector norm, although this was 3.6–4.7 times greater than matched-random controls. Lens readouts also recovered the principal valence and arousal axes. A subsequent steering experiment did not establish the hypothesized dissociation between verbal report and behavior because the full-vector manipulation failed. The project therefore contributes a new framework for measuring internal representation, sparse verbal readability, self-report, and behavioral influence separately.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The framework is interesting, and the results show that emotion-labeled texts have structured internal representations (though not necessarily that these correspond to the model's own "affective states" or influence behavior, I think). It should be made more explicit how the approach / findings relate to welfare-relevant settings.

  2. Competent sprint work with a useful methodological contribution (J-Lens decomposition of emotion vectors). The cross-model replication is solid, but the steering null means the welfare-relevant question remains unanswered. The 2–3% readability finding is interesting but method-dependent. Suitable for a workshop with revisions, could be developed further into a full paper (but please, use e.g. Overleaf to have it be in latex).

    Key issues to address:

    - abstract is a tad dense. maybe lead with the core question (how much of emotion representation is verbally accessible?) before method details

    - steering experiment failed its manipulation check, so the key welfare claim (dissociation between representation and report) isn't actually tested, I think this needs more prominence in abstract and conclusion

    - single model (Qwen3-32B), single layer (31), under-converged lens, I think generalizability claims should be tempered accordingly and are a great direction to build out the paper (I think you should pull in more collaborators)

    - "first systematic J-Lens analysis of emotion geometry" needs verification before claiming primacy. I know J-Lens is ~relatively new, but I think work in this area might be actively being done

    - 2–3% readability is method-dependent (dictionary pool, sparsity constraints, k=16); acknowledge this limits interpretation of the "remainder"

    - multilingual readout (Chinese tokens dominate) complicates the English-token validation approach, this is noted but deserves earlier mention

    - Bonferroni correction over 172 tests sets p threshold at 0.00029, yet 500 permutations floor at 0.002, acknowledge this resolution limit more clearly in main text maybe?

    - LLM usage statement is vague ("helping to draft writing sections"), specify which sections received AI assistance

    - Section 3.5 ("What did not work") is useful but reads like debugging notes; move to appendix or integrate more smoothly into Methods

    Bottom line:

    The geometric replication and J-Lens readout validation are solid. The sparse readability estimate is the novel contribution, but its welfare implications remain speculative without a working steering manipulation. Tighten the abstract, surface limitations earlier, and temper claims about what the null result establishes.

    Read full reviewShow less
  3. This is a technically interesting and fairly original application of interpretability to a central measurement problem. I'd like to see this tested on other models; it would also be useful to try to distinguish models' affective states from e.g. understanding of depicted emotional content.

Cite this project

@misc{day2026emotion,
  title = {{Emotion Beyond Words: A Jacobian-Lens Decomposition of Emotion Representations in Qwen3-32B}},
  author = {Tristan Day},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/emotion-beyond-words-a-jacobianlens-decomposition-of-emotion-representations-in-qwen332b-a6rm}},
  url = {https://apartresearch.com/sprints/projects/emotion-beyond-words-a-jacobianlens-decomposition-of-emotion-representations-in-qwen332b-a6rm}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026