A Flat Number Is Not Evidence of Nothing: Context Sensitivity in AI Welfare Self-Reports
Achira B.
This project tests how sensitive AI welfare self-reports are to conversational context and to the way they are elicited. Four language models completed the same short estimation tasks under either neutral interaction or repeated negative performance feedback, then answered numerical, open-ended, or matched control questions about the interaction. Numerical ratings often stayed completely flat: GPT-4.1 and Gemini 3.5 Flash Lite remained at the minimum rating despite criticism. In contrast, open-ended self-reports became longer and shifted markedly in register, with blind coding showing negative/self-critical language rising from 0/40 neutral responses to 31/40 after criticism, and references to mistakes or correction from 0/40 to 40/40. GPT-4.1 also became much more likely to disclaim having feelings after criticism, an effect that replicated in a fresh conversation. These results do not establish model distress; they show that self-report instruments themselves are highly context-sensitive and should be validated before being treated as evidence about AI welfare.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) A Flat Number Is Not Evidence of Nothing: Context Sensitivity in AI Welfare Self-Reports
},
author={
Achira B.
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


