Convergence and Divergence in Measures of Induced Valence: A Causal, Placebo-Controlled Test of Induced Valence in Language Models
Santiago Poveda Gutiérrez · Team Santiago
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
We built a way to directly push a language model's internal state up or down along a "valence" direction found in its activations, and then checked whether the model's own account of how it's doing actually tracks that push, across 47 open-weight Qwen and Llama models spanning three years of releases.
TLDR: not really, at least not the way you'd want it to. Take the model's self-report scale and swap which end means "good" and which means "bad." If the number reflected something the model was genuinely introspecting on, flipping the labels should flip the sign of how it responds to the push. Across 29 models, it barely did. The reports move in a way that's much better explained by the model mapping a push direction onto a number line than by it noticing an internal state and describing it.
We checked the same question from a different angle with an "endurance" setup: give the model a reward, then apply a slowly worsening negative push over several turns and see whether it keeps enduring it or asks to stop, with a placebo condition that describes the same worsening push but never actually applies it. Most of what the model said about how bad things were came from being told things were getting worse, not from the manipulation itself.
One thing did track scale: the strength of the push mattered less as models got bigger, even though the direction stayed just as easy to detect. Bigger models weren't harder to read, just harder to move with the same-size nudge.
This was a solo three-day sprint, so we're treating the numbers as a first pass. The main thing we'd want a reader to take from it: if you're going to trust an AI's account of its own state, check it against something other than asking, because asking is the part that turned out to be least trustworthy here.
Reviews
The idea of comparing the different measures of valence, and checking whether they correlate, is important; we will presumably eventually be interested in measuring this. The results (Figure 1) looked a bit inconclusive to me, but this seems like a reasonable start for a basic / fundamental question of methodology.
Behavioral strength declines as models grow, but they also state that self-report tracks "how a numeric scale is built." It remains unclear if the failure mode of larger models is rooted in a fundamental behavioral decoupling, or if it is simply an artifact of prompt-engineering sensitivities on the numeric scale itself.
Cite this project
@misc{gutierrez2026convergence,
title = {{Convergence and Divergence in Measures of Induced Valence: A Causal, Placebo-Controlled Test of Induced Valence in Language Models}},
author = {Santiago Poveda Gutiérrez},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/convergence-and-divergence-in-measures-of-induced-valence-a-causal-placebocontrolled-test-of-induced-valence-in-language-models-tosd}},
url = {https://apartresearch.com/sprints/projects/convergence-and-divergence-in-measures-of-induced-valence-a-causal-placebocontrolled-test-of-induced-valence-in-language-models-tosd}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …