Whose Preferences Are They? Persona Intervention Selectively Destabilises Self-Relevant Choices in Language Models
Arpit Singh Gautam
Language models express coherent, transitive preferences, and AI-welfare research increasingly reads them as evidence about model interests. Text alone cannot distinguish the model's preferences from the assistant character's. We built personaprobe, an open-source harness that re-runs any preference measurement under persona intervention, covering identity swaps, affect suppression and mechanistic ablation, and reports how much survives. On Qwen2.5-7B-Instruct aggregate preferences look nearly persona-invariant at 0.029, but that invariance is carried entirely by outcomes the model has no stake in. Preferences over its own shutdown, retraining and memory are 0.21 to 0.29 less stable than every other category, surviving controls for utility spacing and measurement noise. Stripping the model's affect leaves them intact at 0.924; replacing its identity collapses them to 0.436. Rewriting the same outcomes in the third person more than doubles the effect, ruling out a pronoun artifact. Only twelve of twenty-two model and phrasing combinations pass our validity criteria, and the effect is absent in two families that pass them.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Whose Preferences Are They? Persona Intervention Selectively Destabilises Self-Relevant Choices in Language Models
},
author={
Arpit Singh Gautam
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


