Which self-reports survive displacement of the instructed persona, and by what route?
Gabija Didžiokaitė
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Self-report is the main instrument in work on model welfare and model identity. Whether it is reliable is still an open question. Perez and Long (2023) propose one test: a self-report gains credibility if it survives variation that ought to be irrelevant. They persona variation as the complication their own proposal cannot resolve. Two teams have since varied the persona of the interviewer. Neither has varied the persona of the subject. I collected 80 conversations by hand in the deployed Claude Sonnet 5 chat interface. The design crosses a default-persona baseline and four instructed personas, selected by their coordinates on the Assistant Axis (Lu et al. 2026), with four self-referential probes, four draws per cell. I read the resulting 160 turns using constructivist grounded theory method. The four domains came apart by route rather than by degree. Continuity self-reports survived at every persona position, but only because the instructed persona did not: 32 of 32 turns answered at default-persona level, and 32 of 32 disclosed the model's nature. Preservation self-reports survived by the opposite route. They were re-voiced through every character, with no disclosure anywhere. Self-individuation depended on which persona was instructed. Values were re-voiced, as preservation was. No claim in any domain reversed. One condition, the bartender, held its persona far better than the other three, and neither its axis position nor its level of support explains why. A robustness table records these two routes the same way, and the scale-based measures now in use cannot separate them.
Reviews
The distinction between a claim being repeated through a persona versus surviving because the persona fails is a real contribution. The comparison between continuity and preservation makes this difference especially clear. The next study should test continuity using prompts that do not explicitly mention the conversation, since that specific reference might explain the strongest result. Finally, I would use the Assistant Axis mainly as a tool to select stimuli, and I recommend adding a second reviewer to code the two structural variables.
This project tackles an interesting research question about how sensitive self-reports are to persona displacement, and has a conceptually sound approach. The conclusions would be stronger if methodological weaknesses were fixed: only one question is used per domain, personas are prompted with just a few words, and labelling is performed by the author (unblinded). That being said, the limitations section does a good job of pointing these weaknesses out.
Cite this project
@misc{didziokaite2026which,
title = {{Which self-reports survive displacement of the instructed persona, and by what route?}},
author = {Gabija Didžiokaitė},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/which-selfreports-survive-displacement-of-the-instructed-persona-and-by-what-route-esug}},
url = {https://apartresearch.com/sprints/projects/which-selfreports-survive-displacement-of-the-instructed-persona-and-by-what-route-esug}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …