WelfareCheck: Evidence Profiles for Welfare-Relevant Signals in Digital Minds
Eswar Vajja · Team Signal & Boundary
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
We built WelfareCheck to test whether apparent preferences in language models remain consistent when measured in different ways. We ran ten complementary tests across 17 models, covering direct reports, choices, trade-offs, repeated decisions, recovery tasks, and introspection. Several models showed meaningful patterns across multiple checks, but none passed the full evidence chain. Instead of producing one welfare score, WelfareCheck shows where evidence holds, where it breaks down, and what should be tested next.

Reviews
The reporting philosophy is right and I'd like to see it adopted. Distinguishing a check that failed from one never reached from one the design omitted, refusing to average unlike methods into a welfare score, and stopping evaluation when an earlier gate fails rather than letting a later raw pattern rescue it — these are the correct commitments for this field, and the package that implements them is genuinely reusable. Pinned tokenizer revisions and locked requirements put it ahead of nearly everything I read this round.
The difficulty is what Figure 2 actually shows. Reading the heatmap, the modal cell is depth 1. All 17 models stopped at answer mapping or basic validity on M5, 16 of 17 on M4, 15 of 17 on M3 and M6. Across 153 units, zero completed their planned checks. The paper reads this as a map of evidence depth, but a profile in which most cells stall before the main test is mostly reporting that the instrument didn't produce interpretable output — that's a fact about the harness, not about the models. When a method fails basic validity on 17 of 17 models, the honest headline is that the method needs debugging, and the paper says the milder "these outcomes limit what these tests tell us" instead.
I'd point to the scoring rule as a likely cause, and I think it's diagnosable. You score candidate answers by summing log-probability over all tokens of the answer and softmaxing across options. Unnormalized sequence log-probability is monotonically penalised by length, so an option that happens to tokenize longer gets systematically lower probability regardless of meaning. If the allowed answers across your methods differ in token length — and "stay with the current task" versus "switch" almost certainly do — then the mapping check would fail or produce degenerate distributions exactly as observed. Length-normalizing (mean log-prob per token), or scoring a single-token label with the meanings supplied in the prompt, would be the first thing to try. Worth reporting the distribution of π values for failed units too; that would immediately show whether failures look like near-uniform outputs, degenerate one-hot outputs, or parse errors, and the three imply different fixes.
There's also a tension between the paper's stated principle and its main figure. You are emphatic that unlike methods are never placed on a common scale, and then present depth across all ten methods on one 0–5 colour scale. Depth 3 in M7 and depth 3 in M9 are different amounts of different evidence, and a reader scanning rows will compare them anyway — that's what a heatmap is for. Either give each method its own scale with its check count labelled, or show the check names on the axis so a cell's meaning is visible rather than implied.
On Method 10: running activation-injection introspection across 17 models as one method among ten, with no reported detail about what was injected, where, at what strength, or how detection was scored, isn't enough for a reader to evaluate. All 17 passing basic validity and none completing the causal test is consistent with the causal stage never really running. Given how much work this single technique takes to do properly, I'd either give it a proper methods subsection or drop it and say the framework has a slot for it.
The scale is the underlying problem. Ten methods across 17 models in a sprint meant no method got the piloting that would have caught the answer-mapping failures, and the study bought breadth at the cost of any complete result. Two or three methods, piloted until they reliably clear basic validity, would have produced a more convincing demonstration of the same framework — the framework's value is that it can carry a result to completion, and nothing here shows it doing so.
Smaller things. The paper says the full design was fixed before any result was examined, but nothing externally timestamps that; since the pre-specification is what makes the ordered-check discipline meaningful, an OSF registration would convert it from an assertion into a check. The "separate reconstruction" and the "independent audit passed all 16 checks" are load-bearing and unattributed — say who or what performed them. Method names differ between Table 1 and Figure 1 (task ranking versus tournament, delayed choice versus intertemporal). The log-probability formula collides with the sentence around it, so "we compute... and then normalize across the allowed answers:" runs directly into "In plain terms." And the claim that 16 of 17 models reproduced task rankings on unseen tasks is the most interesting positive here, but its controls weren't run, so I'd resist calling it a repeated pattern until they are.
The prose is clear, plain, and well-organized, and the discussion is careful about what the results don't show. That care is why the execution score frustrates me: the framing and the artifacts are ahead of the measurements they carry.
Read full reviewShow less
No info about what the model did, chose, reported, or traded off.
Cite this project
@misc{vajja2026welfarecheck,
title = {{WelfareCheck: Evidence Profiles for Welfare-Relevant Signals in Digital Minds}},
author = {Eswar Vajja},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/welfarecheck-evidence-profiles-for-welfarerelevant-signals-in-digital-minds-3vdm}},
url = {https://apartresearch.com/sprints/projects/welfarecheck-evidence-profiles-for-welfarerelevant-signals-in-digital-minds-3vdm}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …