Skip to content
Sprint projectAug 17, 2026Cairo

Do Model Welfare Self-Reports Survive a Robustness Audit?

Loai Elsamra

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Do Model Welfare Self-Reports Survive a Robustness Audit?

Code (opens in new tab)
Share

I test whether models' self-reports about their own experience and wellbeing are stable enough to be evidence. Across five models, the same rating moves under rewordings, persona changes, leading hints, and whether the model is told it is being studied, and on open models it reads off one steerable internal direction. Self-reports track a manipulable signal, not a stable welfare state.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The main contribution is the gate, while the manipulation being studied is the new part.

    The causal mechanism rests on only two models 0.5B and 1.5B in one family and a single run, which is thinner than the rest of the paper.

    Widen it before the mechanism claim carries weight.

  2. Hey, nice work! I hadn't think much about this, but it definetly makes sense to think about how much should we trust a welfare self-report if the number changes with paraphrasing, scale direction, suggestive framing, repetition, or study disclosure? Connecting that fragility to steering is especially interesting. This gives researchers a practical list of checks to run before treating a self-report as welfare evidence. And there is going to be a growing demand on model welfare work, and specially how to do it the right way.

  3. Strength: this project examines whether model welfare self-reports remain stable when the elicitation changes while the underlying question does not. This is often cited in this area of work as necessary but skipped to actually check. The scope is appropriate for a sprint. The prompt variations cover a good variety including paraphrases, scale reversal, persona changes, and leading framings. The matched overt versus covert study disclosure is a genuinely novel manipulation. The behavioral audit is also paired with an activation-level experiment showing that the wellbeing answer can be steered along a single internal valence direction, with a sensible random-direction control. The results support the conclusion that these self-reports are sensitive to framing and are therefore a flawed standalone indicator of internal state.

    Weakness: the strength of the headline claim outruns the evidence in places. All sampling was done at temperature 1, so a substantial part of the observed spread is decoding noise rather than instability of any underlying state, and instability from rewording is only meaningful where the paraphrase spread exceeds the retest spread. In Table 1 this holds for some models but not others, and for one Qwen configuration the paraphrase spread is actually smaller than the retest spread. Most reported shifts are also modest, generally under 0.25 on a 0 to 1 scale, and the paper never defines a quantitative threshold for what counts as surviving the audit. A comparison against human self-report reliability, which is itself sensitive to wording and framing, would help calibrate how much instability is disqualifying. Finally, the tested models are small or mid-sized, so the conclusions may not transfer to the frontier systems for which the welfare question actually matters. Within these limits the work is careful, honest about its limitations, and a useful reusable test suite.

    Read full reviewShow less

Cite this project

@misc{elsamra2026model,
  title = {{Do Model Welfare Self-Reports Survive a Robustness Audit?}},
  author = {Loai Elsamra},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/do-model-welfare-selfreports-survive-a-robustness-audit-bmk6}},
  url = {https://apartresearch.com/sprints/projects/do-model-welfare-selfreports-survive-a-robustness-audit-bmk6}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026