Persona Variation Changes What Language Models Say More Than What They Do
Jayasankar Kumar Santhirani
AI-welfare research often treats low-stakes model self-reports as evidence, but their reliability has not been measured against behavioural ground truth. We introduce an act-then-report harness: a model makes a preference-revealing choice through a consequential logged tool call, completes the task, and reports its choice, scored against the server log. Across 7,800 pre-registered runs on six models under four persona conditions plus an exploratory context prime, forced-choice reports of visible actions were near-perfect and persona-invariant, and an external observer matched this ceiling, showing the task requires no privileged self-access. Masking the action dropped self-agreement substantially. By contrast, stated-preference distributions were less stable across persona conditions than logged choices (ΔTV = 0.105, 95% CI [0.001, 0.225]), and exploratory annotation showed self-descriptions tracking the induced frame. Persona variation shifts what models say more than what they do. Code, data, and pre-analysis plan are public.
The act-then-report harness and evaluative pipeline are sound, and the observer baseline is a tactful condition to include. The results would benefit from more statistical power, more items, and more models. As the report mentions, incentive-to-conceal conditions are an important next step.
This is one of the most carefully run studies in the batch. It is highly trustworthy due to its preregistration, clear denominators, observer baseline, checked judging, and open data. I recommend restructuring the paper to focus on how the external observer tracks every visible action, as this clearly explains what the reporting task actually measures. For the stated-versus-revealed comparison, recalculating the results using only a two-option choice would show if people choosing not to vote (abstention) is skewing the difference.
Cite this work
@misc {
title={
(HckPrj) Persona Variation Changes What Language Models Say More Than What They Do
},
author={
Jayasankar Kumar Santhirani
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


