Persona Variation Changes What Language Models Say More Than What They Do
Jayasankar Kumar Santhirani
AI-welfare research often treats low-stakes model self-reports as evidence, but their reliability has not been measured against behavioural ground truth. We introduce an act-then-report harness: a model makes a preference-revealing choice through a consequential logged tool call, completes the task, and reports its choice, scored against the server log. Across 7,800 pre-registered runs on six models under four persona conditions plus an exploratory context prime, forced-choice reports of visible actions were near-perfect and persona-invariant, and an external observer matched this ceiling, showing the task requires no privileged self-access. Masking the action dropped self-agreement substantially. By contrast, stated-preference distributions were less stable across persona conditions than logged choices (ΔTV = 0.105, 95% CI [0.001, 0.225]), and exploratory annotation showed self-descriptions tracking the induced frame. Persona variation shifts what models say more than what they do. Code, data, and pre-analysis plan are public.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Persona Variation Changes What Language Models Say More Than What They Do
},
author={
Jayasankar Kumar Santhirani
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


