Base Model Persona Inference can Predict Misalignment and Surface Agreement
Arush Tagade, Taslim Mahbub · Team Praxis Research
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Language models have a consistent assistant persona that is the product of post-training but this assistant character is not well understood. We introduce base model persona inference that assigns a persona to a post-trained LLM response and show that continuations sampled from this persona can: 1) predict broad misalignment caused by harmful responses and 2) provide consistent frames of analysis between revealed and stated scenario preferences. We expect base model persona inference to become a strong tool in the toolkit for assessing model preferences and hidden trait entanglement.
Reviews
The submission focuses on an interesting problem, whether we can use base models to make predictions about post-trained behaviour. The authors had some good ideas for experiments, although the fact that the authors choose a specific persona undermines conclusions. For the author’s conclusions to be supported, I would want to see the predictive power of base models beyond post-trained models.
The central idea in this work is that identifying the persona responsible for some model output will help us to predict future behaviour. While this intuition seems sound to me, there are several aspects of the work that I struggle to understand. In particular: (i) What's the advantage of using base models, specifically, to infer personas from posttrained model outputs? (ii) What exactly is the sense in which persona inferences predict misalignment - in which contexts should we expect misaligned behaviour? (iii) What is the finding about stated and revealed preferences? The work could be improved by providing clearer answers to these questions.
Cite this project
@misc{tagade2026base,
title = {{Base Model Persona Inference can Predict Misalignment and Surface Agreement}},
author = {Arush Tagade and Taslim Mahbub},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/base-model-persona-inference-can-predict-misalignment-and-surface-agreement-ua8o}},
url = {https://apartresearch.com/sprints/projects/base-model-persona-inference-can-predict-misalignment-and-surface-agreement-ua8o}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …