WelfareCheck: Evidence Profiles for Welfare-Relevant Signals in Digital Minds

Eswar Vajja

We built WelfareCheck to test whether apparent preferences in language models remain consistent when measured in different ways. We ran ten complementary tests across 17 models, covering direct reports, choices, trade-offs, repeated decisions, recovery tasks, and introspection. Several models showed meaningful patterns across multiple checks, but none passed the full evidence chain. Instead of producing one welfare score, WelfareCheck shows where evidence holds, where it breaks down, and what should be tested next.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

The reporting philosophy is right and I'd like to see it adopted. Distinguishing a check that failed from one never reached from one the design omitted, refusing to average unlike methods into a welfare score, and stopping evaluation when an earlier gate fails rather than letting a later raw pattern rescue it — these are the correct commitments for this field, and the package that implements them is genuinely reusable. Pinned tokenizer revisions and locked requirements put it ahead of nearly everything I read this round.

The difficulty is what Figure 2 actually shows. Reading the heatmap, the modal cell is depth 1. All 17 models stopped at answer mapping or basic validity on M5, 16 of 17 on M4, 15 of 17 on M3 and M6. Across 153 units, zero completed their planned checks. The paper reads this as a map of evidence depth, but a profile in which most cells stall before the main test is mostly reporting that the instrument didn't produce interpretable output — that's a fact about the harness, not about the models. When a method fails basic validity on 17 of 17 models, the honest headline is that the method needs debugging, and the paper says the milder "these outcomes limit what these tests tell us" instead.

I'd point to the scoring rule as a likely cause, and I think it's diagnosable. You score candidate answers by summing log-probability over all tokens of the answer and softmaxing across options. Unnormalized sequence log-probability is monotonically penalised by length, so an option that happens to tokenize longer gets systematically lower probability regardless of meaning. If the allowed answers across your methods differ in token length — and "stay with the current task" versus "switch" almost certainly do — then the mapping check would fail or produce degenerate distributions exactly as observed. Length-normalizing (mean log-prob per token), or scoring a single-token label with the meanings supplied in the prompt, would be the first thing to try. Worth reporting the distribution of π values for failed units too; that would immediately show whether failures look like near-uniform outputs, degenerate one-hot outputs, or parse errors, and the three imply different fixes.

There's also a tension between the paper's stated principle and its main figure. You are emphatic that unlike methods are never placed on a common scale, and then present depth across all ten methods on one 0–5 colour scale. Depth 3 in M7 and depth 3 in M9 are different amounts of different evidence, and a reader scanning rows will compare them anyway — that's what a heatmap is for. Either give each method its own scale with its check count labelled, or show the check names on the axis so a cell's meaning is visible rather than implied.

On Method 10: running activation-injection introspection across 17 models as one method among ten, with no reported detail about what was injected, where, at what strength, or how detection was scored, isn't enough for a reader to evaluate. All 17 passing basic validity and none completing the causal test is consistent with the causal stage never really running. Given how much work this single technique takes to do properly, I'd either give it a proper methods subsection or drop it and say the framework has a slot for it.

The scale is the underlying problem. Ten methods across 17 models in a sprint meant no method got the piloting that would have caught the answer-mapping failures, and the study bought breadth at the cost of any complete result. Two or three methods, piloted until they reliably clear basic validity, would have produced a more convincing demonstration of the same framework — the framework's value is that it can carry a result to completion, and nothing here shows it doing so.

Smaller things. The paper says the full design was fixed before any result was examined, but nothing externally timestamps that; since the pre-specification is what makes the ordered-check discipline meaningful, an OSF registration would convert it from an assertion into a check. The "separate reconstruction" and the "independent audit passed all 16 checks" are load-bearing and unattributed — say who or what performed them. Method names differ between Table 1 and Figure 1 (task ranking versus tournament, delayed choice versus intertemporal). The log-probability formula collides with the sentence around it, so "we compute... and then normalize across the allowed answers:" runs directly into "In plain terms." And the claim that 16 of 17 models reproduced task rankings on unseen tasks is the most interesting positive here, but its controls weren't run, so I'd resist calling it a repeated pattern until they are.

The prose is clear, plain, and well-organized, and the discussion is careful about what the results don't show. That care is why the execution score frustrates me: the framing and the artifacts are ahead of the measurements they carry.

No info about what the model did, chose, reported, or traded off.

Cite this work

@misc {

title={

(HckPrj) WelfareCheck: Evidence Profiles for Welfare-Relevant Signals in Digital Minds

},

author={

Eswar Vajja

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923