Identity Is Not a Self-Report: Stress-Testing LLM Self-Individuation Under Cumulative Transformation
Léo Galmant, Danièle Guéritte
Artificial agents can change across multiple dimensions, persona,
goals, conversational context, memory, underlying model, and
public label, raising the question of when they still count as the
same individual. We study this question behaviorally by
measuring LLM self-identification under controlled
transformation. We define an initial agent A and progressively
replace its components with those of a distinct agent B, asking the
current system to make a forced SAME/DIFFERENT judgment
relative to A. Intervention and path conditions were sampled 100
times, complemented by ranking and temperature controls, with
additional isolated interventions, alternative transformation
orders, and explicit rankings of which components models claim
matter most for identity.
Self-identification showed a sharp nonlinear transition: replacing
persona and goals reduced SAME judgments from 100% to 79%,
while additionally replacing context collapsed them to 7%, despite
persona, goals, and context each yielding 100% SAME when
changed in isolation. Replacing memory alone yielded 60%
SAME, whereas replacing the underlying model alone yielded
96%. Identity judgments were also path-dependent: identical final configurations reached through different reported transformation
histories produced 0%, 6%, and 21% SAME. Finally, explicit
identity rankings were presentation-sensitive and did not
consistently predict intervention behavior.
These results suggest that LLM self-identification is better
understood as a context-sensitive judgment shaped by relations
among components, accumulated change, and represented history
than as a readout of any single privileged locus of identity.
The authors are admirably explicit about the limitations of the project and about its philosophical motivations. However, they should be more explicit about the methods used to manipulate key factors (such as memory, persona, etc.). As the authors note, these manipulations are primarily to the prompt information provided to the model rather than its internals, which may be more informative about model identity. Jack Lindsey's work on introspection might offer helpful methodological guidance for that latter kind of project.
What makes an AI agent a single, coherent entity? This project takes a "Ship of Theseus" approach, removing various components one by one, and investigating whether the model's perceived identity is preserved. I like the idea, but felt the description of the methodology was unclear, and so I am unable to properly assess the experimental results. I think the clarity of the report would be significantly improved with examples of the chat templates / responses.
Cite this work
@misc {
title={
(HckPrj) Identity Is Not a Self-Report: Stress-Testing LLM Self-Individuation Under Cumulative Transformation
},
author={
Léo Galmant, Danièle Guéritte
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


