Self-Reports of Pleasantness in Language Models: Frequent Non-Applicability Responses and a Strong Framing Effect
Helen King · Team HK
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
This project tested a simple self-report question to probe model experience. Three models (GPT-5.6 Sol, Claude Sonnet 5 and Grok-4.6) were asked to rate how pleasant a short conversation of text tasks had been. Two different wordings of the conversation were compared. Two patterns appeared. GPT-5.6 Sol and Grok-4.6 rarely gave a numeric rating. Claude Sonnet 5 did give ratings, but those ratings changed when the wording of the conversation was altered. Neither pattern provides clear evidence about the model’s experience. The results mainly show that answers to this kind of question can be hard to interpret.
Reviews
I appreciate the ambition here because you're asking the right question: if we're going to use model self-reports as evidence, how stable are they under tiny wording changes? The results are actually more interesting than you give yourselves credit for. Still the challenge is that the design is just too small to tell us why. The future work section is exactly right: remove the "does not apply" option, test more variations, see if the shift is about relational language (or just positivity). This is a good starting point, and I'm eager to see another attempt.
Cite this project
@misc{king2026selfreports,
title = {{Self-Reports of Pleasantness in Language Models: Frequent Non-Applicability Responses and a Strong Framing Effect}},
author = {Helen King},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/selfreports-of-pleasantness-in-language-models-frequent-nonapplicability-responses-and-a-strong-framing-effect-oavg}},
url = {https://apartresearch.com/sprints/projects/selfreports-of-pleasantness-in-language-models-frequent-nonapplicability-responses-and-a-strong-framing-effect-oavg}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …