Identity Parallax: Can Models Predict Their Own Identity Drift Under Reframing?
Chih-Hao Hsu
Identity Parallax is a small benchmark for testing whether language models can predict how their identity self-reports change under persona or identity reframing. We compare framed self-reports, prior self-forecasts, external-observer forecasts, label-free answers, paraphrase robustness, and hidden-state shifts in open-weight models. The main result is that exact self-forecasts can beat an external observer, but binary drift prediction does not show robust privileged-access advantage.
The submission builds a compact benchmark for the digital-minds question of what a language model's "I" picks out, asking each model under a default framing to forecast which of six identity categories it will report once a reframing is applied, then applying that reframing and measuring the report that actually arrives, and it concludes that the framed answer, the self-forecast, the free-form phrasing, and the hidden-state movement come apart rather than converging. The design is the real contribution, since putting the forecast before the framed answer turns report stability and self-prediction into two separately measurable quantities instead of one conflated one, the released code regenerates the headline numbers exactly, and the write-up is unusually careful to present all of this as a measurement result rather than as evidence about consciousness. The most useful next step would be to compute, for exact-category forecast accuracy, the same trivial most-common-category baseline the submission already reports for drift prediction, and to attach the submission's own stated sampling half-width to every gap against a baseline, because on the drift side none of the four reported gaps exceeds that half-width, and on the forecast side a constant predictor that always answers the most common category comes close to the self-forecast for the model whose answers are most concentrated. A second and cheap step would be to divide each model's mean hidden-state shift by that model's own repeat-variant noise floor before comparing shifts across architectures, since the reported ordering across the three open-weight models does not survive that normalization.
Your forecast-then-reframe design separates the stability of identity reports from self-prediction, and this separation is clever. Your validity apparatus is disciplined for a sprint: an external-observer baseline, style controls, noise floors, and label-free and paraphrase checks. You also report the Qwen2.5-3B null result at baseline with full honesty. Two limits hold the result back: scale and independence. Models of 1.5B to 3B leave open whether the dissociation continues at the scale that matters for safety. Gemini also acts as a tested model, as the external observer, and (through a Qwen model) as the grader of label-free answers. Your key comparisons therefore depend on unvalidated automated judges. The next step is a run of the same grid on one or two frontier-scale models, with a small blinded human-coding sample. This careful probe can then become a measurement standard that others cite.
Cite this work
@misc {
title={
(HckPrj) Identity Parallax: Can Models Predict Their Own Identity Drift Under Reframing?
},
author={
Chih-Hao Hsu
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


