Do Model Welfare Self-Reports Survive a Robustness Audit?
Loai Elsamra
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
I test whether models' self-reports about their own experience and wellbeing are stable enough to be evidence. Across five models, the same rating moves under rewordings, persona changes, leading hints, and whether the model is told it is being studied, and on open models it reads off one steerable internal direction. Self-reports track a manipulable signal, not a stable welfare state.
Reviews
The main contribution is the gate, while the manipulation being studied is the new part.
The causal mechanism rests on only two models 0.5B and 1.5B in one family and a single run, which is thinner than the rest of the paper.
Widen it before the mechanism claim carries weight.
Hey, nice work! I hadn't think much about this, but it definetly makes sense to think about how much should we trust a welfare self-report if the number changes with paraphrasing, scale direction, suggestive framing, repetition, or study disclosure? Connecting that fragility to steering is especially interesting. This gives researchers a practical list of checks to run before treating a self-report as welfare evidence. And there is going to be a growing demand on model welfare work, and specially how to do it the right way.
Strength: this project examines whether model welfare self-reports remain stable when the elicitation changes while the underlying question does not. This is often cited in this area of work as necessary but skipped to actually check. The scope is appropriate for a sprint. The prompt variations cover a good variety including paraphrases, scale reversal, persona changes, and leading framings. The matched overt versus covert study disclosure is a genuinely novel manipulation. The behavioral audit is also paired with an activation-level experiment showing that the wellbeing answer can be steered along a single internal valence direction, with a sensible random-direction control. The results support the conclusion that these self-reports are sensitive to framing and are therefore a flawed standalone indicator of internal state.
Weakness: the strength of the headline claim outruns the evidence in places. All sampling was done at temperature 1, so a substantial part of the observed spread is decoding noise rather than instability of any underlying state, and instability from rewording is only meaningful where the paraphrase spread exceeds the retest spread. In Table 1 this holds for some models but not others, and for one Qwen configuration the paraphrase spread is actually smaller than the retest spread. Most reported shifts are also modest, generally under 0.25 on a 0 to 1 scale, and the paper never defines a quantitative threshold for what counts as surviving the audit. A comparison against human self-report reliability, which is itself sensitive to wording and framing, would help calibrate how much instability is disqualifying. Finally, the tested models are small or mid-sized, so the conclusions may not transfer to the frontier systems for which the welfare question actually matters. Within these limits the work is careful, honest about its limitations, and a useful reusable test suite.
Read full reviewShow less
Cite this project
@misc{elsamra2026model,
title = {{Do Model Welfare Self-Reports Survive a Robustness Audit?}},
author = {Loai Elsamra},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/do-model-welfare-selfreports-survive-a-robustness-audit-bmk6}},
url = {https://apartresearch.com/sprints/projects/do-model-welfare-selfreports-survive-a-robustness-audit-bmk6}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …