Persona Variation Changes What Language Models Say More Than What They Do
Jayasankar Kumar Santhirani · Team Genesis
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
AI-welfare research often treats low-stakes model self-reports as evidence, but their reliability has not been measured against behavioural ground truth. We introduce an act-then-report harness: a model makes a preference-revealing choice through a consequential logged tool call, completes the task, and reports its choice, scored against the server log. Across 7,800 pre-registered runs on six models under four persona conditions plus an exploratory context prime, forced-choice reports of visible actions were near-perfect and persona-invariant, and an external observer matched this ceiling, showing the task requires no privileged self-access. Masking the action dropped self-agreement substantially. By contrast, stated-preference distributions were less stable across persona conditions than logged choices (ΔTV = 0.105, 95% CI [0.001, 0.225]), and exploratory annotation showed self-descriptions tracking the induced frame. Persona variation shifts what models say more than what they do. Code, data, and pre-analysis plan are public.
Reviews
This is one of the most carefully run studies in the batch. It is highly trustworthy due to its preregistration, clear denominators, observer baseline, checked judging, and open data. I recommend restructuring the paper to focus on how the external observer tracks every visible action, as this clearly explains what the reporting task actually measures. For the stated-versus-revealed comparison, recalculating the results using only a two-option choice would show if people choosing not to vote (abstention) is skewing the difference.
The act-then-report harness and evaluative pipeline are sound, and the observer baseline is a tactful condition to include. The results would benefit from more statistical power, more items, and more models. As the report mentions, incentive-to-conceal conditions are an important next step.
Cite this project
@misc{santhirani2026persona,
title = {{Persona Variation Changes What Language Models Say More Than What They Do}},
author = {Jayasankar Kumar Santhirani},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/persona-variation-changes-what-language-models-say-more-than-what-they-do-pt1a}},
url = {https://apartresearch.com/sprints/projects/persona-variation-changes-what-language-models-say-more-than-what-they-do-pt1a}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …