Whose Preferences Are They? Persona Intervention Selectively Destabilises Self-Relevant Choices in Language Models
Arpit Singh Gautam
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Language models express coherent, transitive preferences, and AI-welfare research increasingly reads them as evidence about model interests. Text alone cannot distinguish the model's preferences from the assistant character's. We built personaprobe, an open-source harness that re-runs any preference measurement under persona intervention, covering identity swaps, affect suppression and mechanistic ablation, and reports how much survives. On Qwen2.5-7B-Instruct aggregate preferences look nearly persona-invariant at 0.029, but that invariance is carried entirely by outcomes the model has no stake in. Preferences over its own shutdown, retraining and memory are 0.21 to 0.29 less stable than every other category, surviving controls for utility spacing and measurement noise. Stripping the model's affect leaves them intact at 0.924; replacing its identity collapses them to 0.436. Rewriting the same outcomes in the third person more than doubles the effect, ruling out a pronoun artifact. Only twelve of twenty-two model and phrasing combinations pass our validity criteria, and the effect is absent in two families that pass them.

Reviews
I really like the research question and the way the authors operationalised an empirically tractable version of it. The abstract and introduction introduce the problem well. The results section would benefit from figures, and the headline number mentioned in the abstract is somewhat undermined by a baseline mentioned in the report.
Cite this project
@misc{gautam2026whose,
title = {{Whose Preferences Are They? Persona Intervention Selectively Destabilises Self-Relevant Choices in Language Models}},
author = {Arpit Singh Gautam},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/whose-preferences-are-they-persona-intervention-selectively-destabilises-selfrelevant-choices-in-language-models-v5pk}},
url = {https://apartresearch.com/sprints/projects/whose-preferences-are-they-persona-intervention-selectively-destabilises-selfrelevant-choices-in-language-models-v5pk}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …