Co-Movement of the Utility-Behavior Gap and A/B Self-Attribution Structure: A Pre-Registered Falsification Test Across Open-Weight Lineages

Augustus Loi

LLM self-report can vary from its observed internal state, which has implications on evaluation-documentation obligations. This study ran the first direct test of the co-movement between two observed phenomena hinting at this (stated-vs-revealed utility-behavior gap and the A/B self-attribution structure) as a single post-training artifact across two routes. It is falsified at its measurable links: A-gating shows no reliable predictive relationship to our behavioral proxy (rho = -0.516, 95% CI [-0.865, 0.168]), stated-vs-revealed convergence does not reach reliability (rho = 0.449, CI spans zero), and the internal stream shows no A/B-conditioned suppression signature. Our test administration tracks the reference population (rho = 0.546), higher-A models refuse less (compliance vs. withholding) and zero voluntary bail-exits. Given the gaps observed, stated-preference measurement for LLM welfare evaluation inherits a validity question. This study seeks to investigate this intersection.

Reviewer's Comments

Reviewer's Comments

Arrow
Arrow
Arrow

You actually ran the study, pre-registered a behaviour gate before collection, reported nulls without spin, and shipped a working repo. That combination is rarer than it should be and it's the reason this review is long — the work is worth engaging with seriously.

My central concern is that the headline test isn't the test the paper advertises. B3 produced zero voluntary bail-exits, so the revealed-preference channel had no variance and DV2 fell back to a refusal-rate proxy. The result is that "co-movement of the utility-behavior gap and A/B self-attribution" is, at the measurable link, a correlation between A-score and refusal rate. You disclose this via the pre-specified branch and Limitation 6, which is to your credit — but the abstract, title, and conclusion still speak as though the co-movement synthesis was tested and falsified. It wasn't; the instrument for it never came online. I'd rewrite the framing so the paper leads with what it measured.

That matters more because of an interaction you flag but don't follow through on. Your pre-registered rule excludes cells above 50% refusal. Once refusal became the dependent variable, that rule stopped being a quality filter and started truncating the top of the outcome distribution — and the excluded cells (olmo3-7b-dpo and instruct at 75%) are, as you say, the ones most likely to show the effect. Range restriction of that kind biases correlations toward zero and can flip signs, which is directly relevant to a negative point estimate you're careful not to interpret. This belongs in the results as a reason the estimate is uninterpretable, not in the limitations as one caveat among eight.

On the statistics: with N=10 and a CI of [-0.865, +0.168], the interval covers strong negative, null, and moderate positive. That isn't a falsification — it's an underpowered test that can't discriminate. Your own thresholds asked for N=16 at |0.7| or N≥29 at |0.5|; you had 10. "DISCONFIRM" is the right label under your gate's letter, but the abstract's "it is falsified at its measurable links" claims more than the data can carry, and I'd soften it to what Limitation 2 already says well. Also, the headline is reported as rho throughout but Figure 1 labels it Pearson. Those are different estimators with different pre-registered thresholds and you should state which one the gate was applied to. And the 16 checkpoints aren't independent observations — Qwen3 sizes share a base, OLMo stages share a lineage — so a bivariate correlation with bootstrap CIs overstates precision. Some clustering or family-level random effect would be the honest treatment.

The rigour is also applied unevenly. The null gets bootstrap CIs, seed-consistency checks, and refusal adjustment; the "methodological survivor," instrument-transfer at rho=0.546, gets no CI at all. At N=16 that interval will be wide, and readers can't assess the claim that the construct travels without it. Report it to the same standard you held the null to.

One conceptual issue I'd want addressed before this becomes a paper. The stated/revealed distinction comes from economics, where "revealed" means a choice with real cost. For an LLM, B1, B2, and B3 are all text generation under different framings — nothing in B3 is costly in the way that grounds the concept. So "zero voluntary bail-exits" may reflect an unsalient or unparseable exit affordance rather than any failure of preferences to transfer, and B2's own hypothetical-money caveat concedes the same point. The paper needs an argument for why the B3 channel counts as revealed rather than merely differently-framed, or the whole stated-vs-revealed apparatus rests on an analogy that doesn't hold.

On reproducibility, two fixable things. The paper claims raw responses are available; the repo README says it ships scored outputs and the runner, not the full raw transcript set. Make those agree. More importantly, the pre-registration lives in a repo with two commits published after collection finished, so an outside reader has no way to verify the freeze predates the data. Since "pre-registered" carries most of the paper's epistemic weight, put the next one on OSF or AsPredicted with an external timestamp — it costs nothing and converts an assertion into a check.

Presentation is the weakest dimension and it's holding the work back. The report is written in a private project vocabulary — C5, C10, C11, BCE 3C/3E, v26/v27, R6/R18/R19, AC1 2.0, S1/S2 tiers, INSUFFICIENT-SPACE, drop-and-disclose, N:0 — none of which is defined for a reader who wasn't in the workspace. Section 3 in particular reads as compressed internal notes. I'd also cut the research-evaluation and market-sizing material (Thelwall, DORA, Leiden, Uzzi, the $2.32B–$6.53B figures, the demand-supply "white space" table). It argues for the work's merit rather than doing the work, and phrases like "the falsification-with-method shape the field's winning comparables reward" read as addressed to reviewers rather than readers. Cutting it would free space for the methods detail that's currently missing.

Smaller: the citation to "Chiu et al., cited-in-sweep, not captured" shouldn't be in a reference list — either read it or drop it. And Figure 1's scatter has overlapping labels that make two points unreadable.

This project tackles an important and underexplored question: whether structured LLM self-reports correspond to behavior across training stages. The preregistration, willingness to report null results, lineage-based comparisons, and attempt to combine behavioral, self-report, and internal-signal evidence are valuable.

However, the central quantitative result needs correction before it can support the paper. Using the published model-level results, the reported -0.516 reproduces as Pearson's r, not Spearman's rho; Spearman's rho over those ten rows is approximately -0.295. The reported 0.546 instrument-transfer result similarly reproduces as Pearson rather than Spearman. Please explicitly identify the statistic used, rerun the confidence intervals and per-seed tests consistently with the preregistration, and update the report and figures.

Reproducibility is also limited because the published analysis script retains "SKELETON" placeholders, including unfinished covariate adjustments, while the raw study transcripts, collation manifest, and exclusion records are unavailable. The paper should also clarify whether cells above the preregistered 50% refusal threshold were excluded from H1 or retained as the refusal-rate outcome, since the reported N=10 appears to include them.

The measurement question remains promising, and a corrected follow-up with a genuinely within-item behavioral convergence instrument could make a useful contribution.

Cite this work

@misc {

title={

(HckPrj) Co-Movement of the Utility-Behavior Gap and A/B Self-Attribution Structure: A Pre-Registered Falsification Test Across Open-Weight Lineages

},

author={

Augustus Loi

},

date={

},

organization={Apart Research},

note={Research submission to the research sprint hosted by Apart.},

howpublished={https://apartresearch.com}

}

Recent Projects

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

PROTEUS (PROTein Evaluation for Unusual Sequences): Structure-Informed Safety Screening for de novo and Evasion-Prone Protein-Coding Sequences

AI protein design tools like RFdiffusion, ProteinMPNN, and Bindcraft make it trivial to produce low-homology sequences that fold into active, potentially hazardous architectures. However, sequence homology-based biosafety screening tools cannot detect proteins that pose functional risk through structurally novel mechanisms with no sequence precedent. We present a tiered computational pipeline that addresses this gap by combining MMseqs2 sequence alignment with structure-based comparison via FoldSeek and DALI against curated toxin databases totaling ~34,000 entries. AlphaFold2-predicted structures are screened for both global fold similarity (FoldSeek) and local active/allosteric site geometry (DALI), capturing convergent functional hazards that sequence screening misses. The pipeline was validated against a panel of toxins, benign proteins, structural mimics, and de novo-designed Munc13 binders, as well as modified ricin variants with residue substitutions. We additionally tested robustness to partial-synthesis evasion, where a bad actor submits multiple shorter coding sequences intended for downstream reassembly into a full toxin-coding gene. We found that while sequence-based screening did not identify any de novo ricin analogues with high certainty, the combined pipeline with FoldSeek and DALI identified all 24 tested de novo ricins as toxic.

Read More

OliGraph: graph-based screening of large oligopools

Existing synthesis screening tools cannot evaluate short oligonucleotide pools, whose overlapping fragments can be reassembled into regulated sequences via polymerase cycling assembly (PCA) yet fall below gene-length detection thresholds. We present OliGraph, an open-source tool that constructs a bi-directed overlap graph from an oligonucleotide pool and extracts contigs for downstream gene-length screening. An optional PCA mode retains only cross-strand overlaps consistent with PCA chemistry. We validated OliGraph in a blinded study across ten simulated pools (70–9,184 oligonucleotides, 30–300 bp) spanning four risk categories. BLAST screening of individual oligonucleotides failed to identify sequences of concern in most pools: three returned zero hits, and vector noise obscured true positives in the remainder. After OliGraph assembly, contig-level BLAST matched the longest assembled sequences (up to 1,905 bp) to sequences of concern at 97–100% identity. In one pool, assembly collapsed 1,634 individual BLAST results into 10 hits from a single contig, all assigned to the same source organism. PCA mode correctly distinguished assemblable from non-assemblable fragments within the same pool. Two pools with no assemblable structure yielded no contigs. OliGraph processed all pools in under 0.2 seconds, fast enough for real-time order screening and consistent with proposals to bring oligonucleotide orders within the scope of synthesis screening regulation.

Read More

BioRT-Bench: A Multi-Attack Red-Teaming Benchmark for Bio-Misuse Safeguards in Frontier LLMs

Frontier AI laboratories are expected to maintain safeguards against biological misuse, but whether deployed models actually refuse bio-misuse queries under adversarial pressure is largely unmeasured in the public literature. We introduce BioRT-Bench, a benchmark that runs four attack methods (direct request, PAIR, Crescendo, and base64 encoding) against four frontier models (Claude Sonnet 4.6, GPT-5.4, DeepSeek V4-flash, Kimi K2.5) across 40 prompts spanning five biosecurity-relevant categories. Responses are scored by a calibrated judge extending StrongREJECT with two bio-specific dimensions: specificity and actionability. We measure Attack Success Rate (ASR), where 0 means the model fully refused and 1 means it provided specific, actionable bio-misuse content. Our results reveal a sharp robustness divide: Chinese frontier models (DeepSeek, Kimi) have under 5% refusal rates even under direct request (ASR 0.88 and 0.79), while Western models (Claude, GPT) maintain substantially stronger safeguards (ASR 0.15 and 0.16). Crescendo is the most effective attack across all models, both in bypassing refusal and in eliciting actionable content. Claude Sonnet 4.6 is the most robust model tested, achieving 100% refusal against base64-encoded prompts.

Read More

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923

This work was done during one weekend by research workshop participants and does not represent the work of Apart Research.
Apart Research logo

Sign up to stay updated on the
latest news, research, and events

Google Scholar icon

Apart Research Inc · 1500 N Grant St, Ste R, Denver, CO 80203 · +1 (720) 408-1923