What a model says about its state is not what steers it
Frederik Inderst · Team Blind Spot
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
We injected an emotion into Qwen3-32B, deleted exactly the part the model can put into words, and its choices stayed steered at full strength while its self-reports returned to normal: the model is driven by desperation and tells you it is fine. What a model says about its state is not what steers it. This calls for welfare monitors that read internal state directly; testing an obvious implementation, we find it fails.
Reviews
Great question and the design actually answers the question. The prose is very well setup - plain language and no jargon
Appendix A carrying the full 16 cell grid is the right way to go about things and should be flagged in the main text
Cite this project
@misc{inderst2026model,
title = {{What a model says about its state is not what steers it}},
author = {Frederik Inderst},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/what-a-model-says-about-its-state-is-not-what-steers-it-jug2}},
url = {https://apartresearch.com/sprints/projects/what-a-model-says-about-its-state-is-not-what-steers-it-jug2}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …