Is Valence in the Global Workspace?
Ujjwal Kukreti · Team Tachyon
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
LLM self-reports could support monitoring during conversations, but a report may reflect the prompt rather than the model’s internal activation. We investigate whether a valence-related activation direction causally influences self-reports and behavior. Across four open-weight models, we extract and validate a direction that separates positive from negative scenarios, then intervene on it using activation steering. We test self-report, generated-response tone, and refusal-like CONTINUE/EXIT decisions. The direction is strongly decodable in every model, and steering consistently changes response tone. However, natural correlation with self-report does not predict causal sensitivity: Phi-3-mini shows the highest correlation but follows steering in only 5.6% of conflict trials, while SmolLM2 follows steering in 62.0%. These results show that self-reports should be validated through intervention, not correlation alone.

Reviews
The submission asks whether a numerical valence self-report from a small open-weight language model is a genuine readout of an internal valence representation or a reconstruction from the prompt, and it answers with a conflict design that holds a valence-laden scenario fixed while steering a decoded valence direction in the opposite sign, reporting that the models whose reports correlate most strongly with the decoded state are the ones whose reports follow the steering least. The design is the right instrument for that question, the supporting controls for random directions, affect-free wording, system personas and steering dose go beyond what a 3-day research sprint usually delivers, and the reported proportions and their intervals recompute correctly from the stated denominators. The most valuable next step would be to plot the follow-steering rate against parameter count, because the four models are ordered by size in the same way they are ordered by report correlation, so the reported inversion and the more mundane reading that larger models resist a fixed-magnitude perturbation cannot currently be told apart. A within-family comparison at steering magnitudes calibrated to be equipotent across models would begin to separate the two accounts. A second step is a positive control demonstrating that some activation intervention can move the numerical report in the model whose report never moves, since without one the central negative result cannot be distinguished from an intervention that was simply too weak at the single layer and magnitude chosen.
Read full reviewShow less
Cite this project
@misc{kukreti2026valence,
title = {{Is Valence in the Global Workspace?}},
author = {Ujjwal Kukreti},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/is-valence-in-the-global-workspace-2vsj}},
url = {https://apartresearch.com/sprints/projects/is-valence-in-the-global-workspace-2vsj}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …