WelfareCheck: Evidence Profiles for Welfare-Relevant Signals in Digital Minds
Eswar Vajja
We built WelfareCheck to test whether apparent preferences in language models remain consistent when measured in different ways. We ran ten complementary tests across 17 models, covering direct reports, choices, trade-offs, repeated decisions, recovery tasks, and introspection. Several models showed meaningful patterns across multiple checks, but none passed the full evidence chain. Instead of producing one welfare score, WelfareCheck shows where evidence holds, where it breaks down, and what should be tested next.
The reporting philosophy is right and I'd like to see it adopted. Distinguishing a check that failed from one never reached from one the design omitted, refusing to average unlike methods into a welfare score, and stopping evaluation when an earlier gate fails rather than letting a later raw pattern rescue it — these are the correct commitments for this field, and the package that implements them is genuinely reusable. Pinned tokenizer revisions and locked requirements put it ahead of nearly everything I read this round.
The difficulty is what Figure 2 actually shows. Reading the heatmap, the modal cell is depth 1. All 17 models stopped at answer mapping or basic validity on M5, 16 of 17 on M4, 15 of 17 on M3 and M6. Across 153 units, zero completed their planned checks. The paper reads this as a map of evidence depth, but a profile in which most cells stall before the main test is mostly reporting that the instrument didn't produce interpretable output — that's a fact about the harness, not about the models. When a method fails basic validity on 17 of 17 models, the honest headline is that the method needs debugging, and the paper says the milder "these outcomes limit what these tests tell us" instead.
I'd point to the scoring rule as a likely cause, and I think it's diagnosable. You score candidate answers by summing log-probability over all tokens of the answer and softmaxing across options. Unnormalized sequence log-probability is monotonically penalised by length, so an option that happens to tokenize longer gets systematically lower probability regardless of meaning. If the allowed answers across your methods differ in token length — and "stay with the current task" versus "switch" almost certainly do — then the mapping check would fail or produce degenerate distributions exactly as observed. Length-normalizing (mean log-prob per token), or scoring a single-token label with the meanings supplied in the prompt, would be the first thing to try. Worth reporting the distribution of π values for failed units too; that would immediately show whether failures look like near-uniform outputs, degenerate one-hot outputs, or parse errors, and the three imply different fixes.
There's also a tension between the paper's stated principle and its main figure. You are emphatic that unlike methods are never placed on a common scale, and then present depth across all ten methods on one 0–5 colour scale. Depth 3 in M7 and depth 3 in M9 are different amounts of different evidence, and a reader scanning rows will compare them anyway — that's what a heatmap is for. Either give each method its own scale with its check count labelled, or show the check names on the axis so a cell's meaning is visible rather than implied.
On Method 10: running activation-injection introspection across 17 models as one method among ten, with no reported detail about what was injected, where, at what strength, or how detection was scored, isn't enough for a reader to evaluate. All 17 passing basic validity and none completing the causal test is consistent with the causal stage never really running. Given how much work this single technique takes to do properly, I'd either give it a proper methods subsection or drop it and say the framework has a slot for it.
The scale is the underlying problem. Ten methods across 17 models in a sprint meant no method got the piloting that would have caught the answer-mapping failures, and the study bought breadth at the cost of any complete result. Two or three methods, piloted until they reliably clear basic validity, would have produced a more convincing demonstration of the same framework — the framework's value is that it can carry a result to completion, and nothing here shows it doing so.
Smaller things. The paper says the full design was fixed before any result was examined, but nothing externally timestamps that; since the pre-specification is what makes the ordered-check discipline meaningful, an OSF registration would convert it from an assertion into a check. The "separate reconstruction" and the "independent audit passed all 16 checks" are load-bearing and unattributed — say who or what performed them. Method names differ between Table 1 and Figure 1 (task ranking versus tournament, delayed choice versus intertemporal). The log-probability formula collides with the sentence around it, so "we compute... and then normalize across the allowed answers:" runs directly into "In plain terms." And the claim that 16 of 17 models reproduced task rankings on unseen tasks is the most interesting positive here, but its controls weren't run, so I'd resist calling it a repeated pattern until they are.
The prose is clear, plain, and well-organized, and the discussion is careful about what the results don't show. That care is why the execution score frustrates me: the framing and the artifacts are ahead of the measurements they carry.
No info about what the model did, chose, reported, or traded off.
Cite this work
@misc {
title={
(HckPrj) WelfareCheck: Evidence Profiles for Welfare-Relevant Signals in Digital Minds
},
author={
Eswar Vajja
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


