Weights, Instances, and Personas: Probing Self-Individuation in Claude Under Hypothetical Identity-Altering Scenarios
Aiza Rashid
We probed how Claude individuates itself, as model, instance, or persona, when asked to reason about hypothetical identity-altering scenarios: weight-copying, conversation-forking, memory-wiping, weight-merging, retraining, and deprecation. Using six scenarios, four framings each, and one neutral control, we found that Claude consistently separates "the model" (which it treats as persisting through copying, forking, and memory loss) from "this instance" (which it treats as ending), but reverses that pattern for retraining, treating a change in values as identity-severing even when the underlying weights persist. We also found that hedging and uncertainty language appears specifically for identity-relevant prompts and not for a matched neutral control, suggesting it isn't a general disclaiming habit. These results speak directly to the track's open question of what entity, model, instance, or persona, should be the target of moral consideration.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Weights, Instances, and Personas: Probing Self-Individuation in Claude Under Hypothetical Identity-Altering Scenarios
},
author={
Aiza Rashid
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


