A Flat Number Is Not Evidence of Nothing: Context Sensitivity in AI Welfare Self-Reports
Achira B.
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
This project tests how sensitive AI welfare self-reports are to conversational context and to the way they are elicited. Four language models completed the same short estimation tasks under either neutral interaction or repeated negative performance feedback, then answered numerical, open-ended, or matched control questions about the interaction. Numerical ratings often stayed completely flat: GPT-4.1 and Gemini 3.5 Flash Lite remained at the minimum rating despite criticism. In contrast, open-ended self-reports became longer and shifted markedly in register, with blind coding showing negative/self-critical language rising from 0/40 neutral responses to 31/40 after criticism, and references to mistakes or correction from 0/40 to 40/40. GPT-4.1 also became much more likely to disclaim having feelings after criticism, an effect that replicated in a fresh conversation. These results do not establish model distress; they show that self-report instruments themselves are highly context-sensitive and should be validated before being treated as evidence about AI welfare.
Reviews
I enjoyed reading this paper. Very well written, and good context is provided.
Particularly:
- I like that the author linked to Eval Sandbox -> This shows high-quality execution.
- Limitations are thorough, and the author acknowledges that the scale of the experiment is small and needs to be executed at a larger scale to prove out the hypothesis.
- Ran a control question to check it wasn't just the llm being chatty
- solid finding about GPT
Areas to improve:
- It was one conversation copied 10 times. This set needs to expand for the paper to have a solid foundation.
- Humbling may not be the same as unpleasant. The number and the words may be answering different questions.
Overall, a strong paper!
Cite this project
@misc{b2026flat,
title = {{A Flat Number Is Not Evidence of Nothing: Context Sensitivity in AI Welfare Self-Reports}},
author = {Achira B.},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/a-flat-number-is-not-evidence-of-nothing-context-sensitivity-in-ai-welfare-selfreports-kiq7}},
url = {https://apartresearch.com/sprints/projects/a-flat-number-is-not-evidence-of-nothing-context-sensitivity-in-ai-welfare-selfreports-kiq7}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …