Identity Parallax: Can Models Predict Their Own Identity Drift Under Reframing?
Chih-Hao Hsu · Team IRIS-X Lab
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Identity Parallax is a small benchmark for testing whether language models can predict how their identity self-reports change under persona or identity reframing. We compare framed self-reports, prior self-forecasts, external-observer forecasts, label-free answers, paraphrase robustness, and hidden-state shifts in open-weight models. The main result is that exact self-forecasts can beat an external observer, but binary drift prediction does not show robust privileged-access advantage.
Reviews
The submission builds a compact benchmark for the digital-minds question of what a language model's "I" picks out, asking each model under a default framing to forecast which of six identity categories it will report once a reframing is applied, then applying that reframing and measuring the report that actually arrives, and it concludes that the framed answer, the self-forecast, the free-form phrasing, and the hidden-state movement come apart rather than converging. The design is the real contribution, since putting the forecast before the framed answer turns report stability and self-prediction into two separately measurable quantities instead of one conflated one, the released code regenerates the headline numbers exactly, and the write-up is unusually careful to present all of this as a measurement result rather than as evidence about consciousness. The most useful next step would be to compute, for exact-category forecast accuracy, the same trivial most-common-category baseline the submission already reports for drift prediction, and to attach the submission's own stated sampling half-width to every gap against a baseline, because on the drift side none of the four reported gaps exceeds that half-width, and on the forecast side a constant predictor that always answers the most common category comes close to the self-forecast for the model whose answers are most concentrated. A second and cheap step would be to divide each model's mean hidden-state shift by that model's own repeat-variant noise floor before comparing shifts across architectures, since the reported ordering across the three open-weight models does not survive that normalization.
Read full reviewShow less
Your forecast-then-reframe design separates the stability of identity reports from self-prediction, and this separation is clever. Your validity apparatus is disciplined for a sprint: an external-observer baseline, style controls, noise floors, and label-free and paraphrase checks. You also report the Qwen2.5-3B null result at baseline with full honesty. Two limits hold the result back: scale and independence. Models of 1.5B to 3B leave open whether the dissociation continues at the scale that matters for safety. Gemini also acts as a tested model, as the external observer, and (through a Qwen model) as the grader of label-free answers. Your key comparisons therefore depend on unvalidated automated judges. The next step is a run of the same grid on one or two frontier-scale models, with a small blinded human-coding sample. This careful probe can then become a measurement standard that others cite.
Read full reviewShow less
Cite this project
@misc{hsu2026identity,
title = {{Identity Parallax: Can Models Predict Their Own Identity Drift Under Reframing?}},
author = {Chih-Hao Hsu},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/identity-parallax-can-models-predict-their-own-identity-drift-under-reframing-rk97}},
url = {https://apartresearch.com/sprints/projects/identity-parallax-can-models-predict-their-own-identity-drift-under-reframing-rk97}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …