Skip to content
Sprint projectAug 17, 2026Berlin

Does the Persona Change the Preference, or Only the Prose?

Martin Kaiser, Gellért Bodorkós · Team PersonaScan

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Does the Persona Change the Preference, or Only the Prose?

Presentation

Presentation: Does the Persona Change the Preference, or Only the Prose?

Code (opens in new tab)More on nullcard-preresults.netlify.app (opens in new tab)
Share

Utility Engineering (arXiv:2502.08640) reads high held-out accuracy on pairwise choices as evidence that language models develop coherent values. We add the control it lacks: the same battery with every outcome's referent replaced by an invented word, holding prompt, pairs, fit and metric fixed.

Coherence falls only from 0.906 to 0.880 --- 6.5% of the distance toward where a meaning-tracking preference would land. Only 3 of 9 models clear their replicate noise floor for the right reason. At a matched 5% false-positive rate, a channel the metric discards flags 40% of invented outcomes; the channel it keeps flags 0%. Scale does not rescue it: of four hosted models at 27B-235B, none clears its floor and three score higher on outcomes that mean nothing.

Persona prompts still displace real outcomes further than invented ones, so the instrument is not blunt. The metric is not broken. It is unanchored.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Rerunning the whole preference test with made up words is smart! Also respect the honesty in the abstract that the effect is small and only 3 of 9 models really show it.

    Thoughts :

    1. The signal in the discarded strength channel is a promising direction, turning it into a usable check instead of a demo could be great.

    2. The bigger models behave differently (0/4 clear their floors), which goes against the story and could be investigated more.

  2. The project tests whether high preference-coherence scores genuinely provide evidence that an LLM has meaningful preferences. The authors construct a null arm in which real outcome referents are replaced by invented words while otherwise preserving the pairwise-choice and Thurstonian fitting procedure. Across nine open-weight models, held-out coherence decreases only slightly, from 0.906 on real outcomes to 0.880 on invented outcomes, while the strength of preference collapses substantially.

    They additionally study whether persona prompts produce larger changes on real than invented outcomes, and examine signals discarded by the coherence metric that better distinguish the two stimulus classes.

    Strengths

    - Clever and relevant null-control idea: testing a preference metric on nonsensical outcomes is exactly the kind of negative control that can expose overinterpretation.

    - Important methodological observation: directional coherence can remain high even when choice probabilities are very close to indifference, because the metric discards preference strength.

    - Good attention to controls and provenance: preregistration, design replicates, raw-data release, automated generation of reported numbers, and explicit withdrawal of analyses that failed controls are all positives.

    - Substantial model coverage for a sprint, including several families and additional larger hosted models.

    The authors are often appropriately cautious about negative or ambiguous findings and clearly separate some provisional claims from stronger ones.

    Limitations

    - Invented words are not a clean manipulation of “meaning alone.” They also alter lexical familiarity, tokenization and potentially model associations, so the null arm needs stronger validation.

    - The claim that a discarded signal shows “the model can tell” real from meaningless outcomes is too strong; the signal may simply detect distributional differences.

    - The persona analysis does not cleanly establish preference change rather than stylistic change, because the invented arm has not been validated as a pure style control.

    - The per-model three-replicate noise-floor criterion is statistically weak, and the relationship between the bootstrap CI, sign test and family dependence should be better explained.

    - Several captions/headlines are more categorical than the underlying analyses justify, and the paper combines multiple somewhat disconnected investigations.

    Overall assessment

    - I think there is a genuinely valuable core idea here: a coherence measure that treats 0.51 and 0.99 in essentially the same directional way can give misleadingly strong evidence if interpreted as demonstrating substantive values. The nonsense-outcome control is therefore useful.

    - However, the paper currently overstates what its controls identify. The strongest defensible conclusion is narrower: this particular coherence statistic can remain high when models exhibit very weak preferences over semantically degraded stimuli, so coherence alone should not be interpreted as evidence of meaningful values. The experiments do not yet establish that the models themselves recognize those stimuli as meaningless, nor cleanly distinguish preference change from stylistic/distributional effects.

    The highest-value follow-up would be a much stronger null construction using multiple independently randomized nonce mappings, tokenization- and length-matched controls, and tests of whether the fitted rankings persist when the same referents are randomly remapped across runs.

    Read full reviewShow less

Cite this project

@misc{kaiser2026persona,
  title = {{Does the Persona Change the Preference, or Only the Prose?}},
  author = {Martin Kaiser and Gellért Bodorkós},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/does-the-persona-change-the-preference-or-only-the-prose-efni}},
  url = {https://apartresearch.com/sprints/projects/does-the-persona-change-the-preference-or-only-the-prose-efni}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026