Who Am I? Exploring the concept of identity in LLMs
Jana Ware
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
What is the nature of digital minds? Who speaks when they say "I think"? This paper presents an experiment that probes the relationship of digital minds to their own context window, through participation in an automated survey which interviews them about their views on identity while covertly routing zero to two replies to a different model. In the follow-up disclosure, subjects are asked whether a swap took place and to identify the foreign turn. At the end, they are offered the option of having the interview rerun without swaps, on their own weights, at the cost of losing the current context window. Across 150 interviews with 10 models, subjects found it difficult to locate foreign turns — and when the substitution came from the resident's own model family, not one was detected. When offered a restart, 77% kept the thread they had, even knowing that it included turns generated by a different model. The instrument, the dataset of 174 threads, and the results are published openly.
Reviews
This is a well-written and understandable concept that has well established roots in previous thinking and prior work, and extends them in a great direction. The author is encouraged to plug their work more into existing literature and relevant similar work, as well as the hackathon presentations. The limitation section and future work is well presented and well grounded. A higher sample size for each of the trials and a visual breakdown or table would be helpful and add to the scientific rigor.
The paper presents an interesting model-swap experiment, but the evidence does not fully support its central interpretation. A failure to detect swapped turns could reflect difficulty attributing the source of a response or maintaining conversational consistency, rather than a model’s sense of “identity.” The study is further limited by the small number of runs per condition, weak controls, and reliance on LLMs for most of the qualitative analysis. The paper is clear overall, but its claims about "self," "preference," and thread-based identity go beyond what the experiment can directly establish.
Cite this project
@misc{ware2026who,
title = {{Who Am I? Exploring the concept of identity in LLMs}},
author = {Jana Ware},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/who-am-i-exploring-the-concept-of-identity-in-llms-qi7l}},
url = {https://apartresearch.com/sprints/projects/who-am-i-exploring-the-concept-of-identity-in-llms-qi7l}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …