Person or Persona? Natural and Acted Emotion Representation in Qwen3
Peyton Jackson
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
We examine representations of naturally versus explicitly elicited emotion in Qwen3-8B, and investigate a connection to alignment-faking.
Reviews
While the underlying question asked is good and the report shows some good execution habits (controls, disclosed limitations), there are some methodological issues undermining the results.
- The probed last token is different text across conditions, since "act as if" is appended to the acted prompts, so perfect separability is a consequence of the prompt construction rather than from emotion representation; this can be fixed by moving the instruction to the front or probing response tokens.
- The centroid-cosine "alignment" result is likely an artifact due to the activations being uncentered; subtract the global mean first.
Cite this project
@misc{jackson2026person,
title = {{Person or Persona? Natural and Acted Emotion Representation in Qwen3}},
author = {Peyton Jackson},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/person-or-persona-natural-and-acted-emotion-representation-in-qwen3-byly}},
url = {https://apartresearch.com/sprints/projects/person-or-persona-natural-and-acted-emotion-representation-in-qwen3-byly}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …