Identity Is Not a Self-Report: Stress-Testing LLM Self-Individuation Under Cumulative Transformation
Léo Galmant, Danièle Guéritte · Team Léo and Danièle
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Artificial agents can change across multiple dimensions, persona, goals, conversational context, memory, underlying model, and public label, raising the question of when they still count as the same individual. We study this question behaviorally by measuring LLM self-identification under controlled transformation. We define an initial agent A and progressively replace its components with those of a distinct agent B, asking the current system to make a forced SAME/DIFFERENT judgment relative to A. Intervention and path conditions were sampled 100 times, complemented by ranking and temperature controls, with additional isolated interventions, alternative transformation orders, and explicit rankings of which components models claim matter most for identity. Self-identification showed a sharp nonlinear transition: replacing persona and goals reduced SAME judgments from 100% to 79%, while additionally replacing context collapsed them to 7%, despite persona, goals, and context each yielding 100% SAME when changed in isolation. Replacing memory alone yielded 60% SAME, whereas replacing the underlying model alone yielded 96%. Identity judgments were also path-dependent: identical final configurations reached through different reported transformation histories produced 0%, 6%, and 21% SAME. Finally, explicit identity rankings were presentation-sensitive and did not consistently predict intervention behavior. These results suggest that LLM self-identification is better understood as a context-sensitive judgment shaped by relations among components, accumulated change, and represented history than as a readout of any single privileged locus of identity.
Reviews
The authors are admirably explicit about the limitations of the project and about its philosophical motivations. However, they should be more explicit about the methods used to manipulate key factors (such as memory, persona, etc.). As the authors note, these manipulations are primarily to the prompt information provided to the model rather than its internals, which may be more informative about model identity. Jack Lindsey's work on introspection might offer helpful methodological guidance for that latter kind of project.
What makes an AI agent a single, coherent entity? This project takes a "Ship of Theseus" approach, removing various components one by one, and investigating whether the model's perceived identity is preserved. I like the idea, but felt the description of the methodology was unclear, and so I am unable to properly assess the experimental results. I think the clarity of the report would be significantly improved with examples of the chat templates / responses.
Cite this project
@misc{galmant2026identity,
title = {{Identity Is Not a Self-Report: Stress-Testing LLM Self-Individuation Under Cumulative Transformation}},
author = {Léo Galmant and Danièle Guéritte},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/identity-is-not-a-selfreport-stresstesting-llm-selfindividuation-under-cumulative-transformation-puzh}},
url = {https://apartresearch.com/sprints/projects/identity-is-not-a-selfreport-stresstesting-llm-selfindividuation-under-cumulative-transformation-puzh}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …