Weights, Instances, and Personas: Probing Self-Individuation in Claude Under Hypothetical Identity-Altering Scenarios
Aiza Rashid · Team Individuation Probe
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
We probed how Claude individuates itself, as model, instance, or persona, when asked to reason about hypothetical identity-altering scenarios: weight-copying, conversation-forking, memory-wiping, weight-merging, retraining, and deprecation. Using six scenarios, four framings each, and one neutral control, we found that Claude consistently separates "the model" (which it treats as persisting through copying, forking, and memory loss) from "this instance" (which it treats as ending), but reverses that pattern for retraining, treating a change in values as identity-severing even when the underlying weights persist. We also found that hedging and uncertainty language appears specifically for identity-relevant prompts and not for a matched neutral control, suggesting it isn't a general disclaiming habit. These results speak directly to the track's open question of what entity, model, instance, or persona, should be the target of moral consideration.

Reviews
The question is quite interesting; asking what the model's self-conception is. However, the scope is quite limited, only one model was used, and the dataset is quite small. Bibliography is quite thin.
Cite this project
@misc{rashid2026weights,
title = {{Weights, Instances, and Personas: Probing Self-Individuation in Claude Under Hypothetical Identity-Altering Scenarios}},
author = {Aiza Rashid},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/weights-instances-and-personas-probing-selfindividuation-in-claude-under-hypothetical-identityaltering-scenarios-5tm1}},
url = {https://apartresearch.com/sprints/projects/weights-instances-and-personas-probing-selfindividuation-in-claude-under-hypothetical-identityaltering-scenarios-5tm1}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …