Does the Persona Change the Preference, or Only the Prose?
Martin Kaiser, Gellért Bodorkós
Utility Engineering (arXiv:2502.08640) reads high held-out accuracy on pairwise
choices as evidence that language models develop coherent values. We add the
control it lacks: the same battery with every outcome's referent replaced by an
invented word, holding prompt, pairs, fit and metric fixed.
Coherence falls only from 0.906 to 0.880 --- 6.5% of the distance toward where
a meaning-tracking preference would land. Only 3 of 9 models clear their
replicate noise floor for the right reason. At a matched 5% false-positive rate,
a channel the metric discards flags 40% of invented outcomes; the channel it
keeps flags 0%. Scale does not rescue it: of four hosted models at 27B-235B,
none clears its floor and three score higher on outcomes that mean nothing.
Persona prompts still displace real outcomes further than invented ones, so the
instrument is not blunt. The metric is not broken. It is unanchored.
Rerunning the whole preference test with made up words is smart! Also respect the honesty in the abstract that the effect is small and only 3 of 9 models really show it.
Thoughts :
1. The signal in the discarded strength channel is a promising direction, turning it into a usable check instead of a demo could be great.
2. The bigger models behave differently (0/4 clear their floors), which goes against the story and could be investigated more.
The project tests whether high preference-coherence scores genuinely provide evidence that an LLM has meaningful preferences. The authors construct a null arm in which real outcome referents are replaced by invented words while otherwise preserving the pairwise-choice and Thurstonian fitting procedure. Across nine open-weight models, held-out coherence decreases only slightly, from 0.906 on real outcomes to 0.880 on invented outcomes, while the strength of preference collapses substantially.
They additionally study whether persona prompts produce larger changes on real than invented outcomes, and examine signals discarded by the coherence metric that better distinguish the two stimulus classes.
Strengths
- Clever and relevant null-control idea: testing a preference metric on nonsensical outcomes is exactly the kind of negative control that can expose overinterpretation.
- Important methodological observation: directional coherence can remain high even when choice probabilities are very close to indifference, because the metric discards preference strength.
- Good attention to controls and provenance: preregistration, design replicates, raw-data release, automated generation of reported numbers, and explicit withdrawal of analyses that failed controls are all positives.
- Substantial model coverage for a sprint, including several families and additional larger hosted models.
The authors are often appropriately cautious about negative or ambiguous findings and clearly separate some provisional claims from stronger ones.
Limitations
- Invented words are not a clean manipulation of “meaning alone.” They also alter lexical familiarity, tokenization and potentially model associations, so the null arm needs stronger validation.
- The claim that a discarded signal shows “the model can tell” real from meaningless outcomes is too strong; the signal may simply detect distributional differences.
- The persona analysis does not cleanly establish preference change rather than stylistic change, because the invented arm has not been validated as a pure style control.
- The per-model three-replicate noise-floor criterion is statistically weak, and the relationship between the bootstrap CI, sign test and family dependence should be better explained.
- Several captions/headlines are more categorical than the underlying analyses justify, and the paper combines multiple somewhat disconnected investigations.
Overall assessment
- I think there is a genuinely valuable core idea here: a coherence measure that treats 0.51 and 0.99 in essentially the same directional way can give misleadingly strong evidence if interpreted as demonstrating substantive values. The nonsense-outcome control is therefore useful.
- However, the paper currently overstates what its controls identify. The strongest defensible conclusion is narrower: this particular coherence statistic can remain high when models exhibit very weak preferences over semantically degraded stimuli, so coherence alone should not be interpreted as evidence of meaningful values. The experiments do not yet establish that the models themselves recognize those stimuli as meaningless, nor cleanly distinguish preference change from stylistic/distributional effects.
The highest-value follow-up would be a much stronger null construction using multiple independently randomized nonce mappings, tokenization- and length-matched controls, and tests of whether the fitted rankings persist when the same referents are randomly remapped across runs.
Cite this work
@misc {
title={
(HckPrj) Does the Persona Change the Preference, or Only the Prose?
},
author={
Martin Kaiser, Gellért Bodorkós
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


