Is Valence in the Global Workspace?
Ujjwal Kukreti
LLM self-reports could support monitoring during conversations, but a report may reflect the prompt rather than the model’s internal activation. We investigate whether a valence-related activation direction causally influences self-reports and behavior. Across four open-weight models, we extract and validate a direction that separates positive from negative scenarios, then intervene on it using activation steering. We test self-report, generated-response tone, and refusal-like CONTINUE/EXIT decisions. The direction is strongly decodable in every model, and steering consistently changes response tone. However, natural correlation with self-report does not predict causal sensitivity: Phi-3-mini shows the highest correlation but follows steering in only 5.6% of conflict trials, while SmolLM2 follows steering in 62.0%. These results show that self-reports should be validated through intervention, not correlation alone.
The submission asks whether a numerical valence self-report from a small open-weight language model is a genuine readout of an internal valence representation or a reconstruction from the prompt, and it answers with a conflict design that holds a valence-laden scenario fixed while steering a decoded valence direction in the opposite sign, reporting that the models whose reports correlate most strongly with the decoded state are the ones whose reports follow the steering least. The design is the right instrument for that question, the supporting controls for random directions, affect-free wording, system personas and steering dose go beyond what a 3-day research sprint usually delivers, and the reported proportions and their intervals recompute correctly from the stated denominators. The most valuable next step would be to plot the follow-steering rate against parameter count, because the four models are ordered by size in the same way they are ordered by report correlation, so the reported inversion and the more mundane reading that larger models resist a fixed-magnitude perturbation cannot currently be told apart. A within-family comparison at steering magnitudes calibrated to be equipotent across models would begin to separate the two accounts. A second step is a positive control demonstrating that some activation intervention can move the numerical report in the model whose report never moves, since without one the central negative result cannot be distinguished from an intervention that was simply too weak at the single layer and magnitude chosen.
Cite this work
@misc {
title={
(HckPrj) Is Valence in the Global Workspace?
},
author={
Ujjwal Kukreti
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


