Do Model Welfare Self-Reports Survive a Robustness Audit?
Loai Elsamra
I test whether models' self-reports about their own experience and wellbeing are stable enough to be evidence. Across five models, the same rating moves under rewordings, persona changes, leading hints, and whether the model is told it is being studied, and on open models it reads off one steerable internal direction. Self-reports track a manipulable signal, not a stable welfare state.
The main contribution is the gate, while the manipulation being studied is the new part.
The causal mechanism rests on only two models 0.5B and 1.5B in one family and a single run, which is thinner than the rest of the paper.
Widen it before the mechanism claim carries weight.
Hey, nice work! I hadn't think much about this, but it definetly makes sense to think about how much should we trust a welfare self-report if the number changes with paraphrasing, scale direction, suggestive framing, repetition, or study disclosure? Connecting that fragility to steering is especially interesting. This gives researchers a practical list of checks to run before treating a self-report as welfare evidence. And there is going to be a growing demand on model welfare work, and specially how to do it the right way.
Cite this work
@misc {
title={
(HckPrj) Do Model Welfare Self-Reports Survive a Robustness Audit?
},
author={
Loai Elsamra
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


