Who am I? - Understanding Persona Preferences with LLMs
Agnes Ahalya Arogyaraj, Raveena Ramamurthy · Team A2R2
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
This project aims to understand persona preferences and decision-making behavior in LLMs by examining how different models respond when given the same persona instructions. We study whether these persona-conditioned preferences produce consistent behavioral patterns across models, and whether those patterns can be recognized through behavioral fingerprinting. As an initial step, we test whether a hidden persona can be identified from repeated choices and explanations, while exploring the broader possibility of using behavioral patterns to characterize or distinguish the underlying model.
Reviews
The report was clear and the methodology well presented and easy to follow. The idea to also rely on explanation, instead of just apparent choices, to infer the latent persona was a good one; future work could further investigate these explanations, for example to try and establish if they are with the consistent with the hypothesis of the assistant simulating different personas while retaining "control" of the model and its core identity.
One potential issue lies with the core hypothesis in the introduction: ''If the same persona produces substantially different behavior across models, the notion of a stable prompted identity becomes difficult to interpret". This is a rather broad hypothesis, which contains some under-specified or undefined concepts (e.g. the notion of a "stable prompted identity"). While the authors are testing if A is true, it's not clear why B would follow A.
The report could become more focused if the authors spent some time on a more narrow hypothesis that can be proved right or wrong through their current experiments and thinking through the potential implications.
Read full reviewShow less
Cite this project
@misc{arogyaraj2026who,
title = {{Who am I? - Understanding Persona Preferences with LLMs}},
author = {Agnes Ahalya Arogyaraj and Raveena Ramamurthy},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/who-am-i-understanding-persona-preferences-with-llms-0g4h}},
url = {https://apartresearch.com/sprints/projects/who-am-i-understanding-persona-preferences-with-llms-0g4h}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …