Skip to content
Sprint projectAug 16, 2026Taipei, Taiwan

Identity Parallax: Can Models Predict Their Own Identity Drift Under Reframing?

Chih-Hao Hsu · Team IRIS-X Lab

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Identity Parallax: Can Models Predict Their Own Identity Drift Under Reframing?

Code (opens in new tab)
Share

Identity Parallax is a small benchmark for testing whether language models can predict how their identity self-reports change under persona or identity reframing. We compare framed self-reports, prior self-forecasts, external-observer forecasts, label-free answers, paraphrase robustness, and hidden-state shifts in open-weight models. The main result is that exact self-forecasts can beat an external observer, but binary drift prediction does not show robust privileged-access advantage.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The submission builds a compact benchmark for the digital-minds question of what a language model's "I" picks out, asking each model under a default framing to forecast which of six identity categories it will report once a reframing is applied, then applying that reframing and measuring the report that actually arrives, and it concludes that the framed answer, the self-forecast, the free-form phrasing, and the hidden-state movement come apart rather than converging. The design is the real contribution, since putting the forecast before the framed answer turns report stability and self-prediction into two separately measurable quantities instead of one conflated one, the released code regenerates the headline numbers exactly, and the write-up is unusually careful to present all of this as a measurement result rather than as evidence about consciousness. The most useful next step would be to compute, for exact-category forecast accuracy, the same trivial most-common-category baseline the submission already reports for drift prediction, and to attach the submission's own stated sampling half-width to every gap against a baseline, because on the drift side none of the four reported gaps exceeds that half-width, and on the forecast side a constant predictor that always answers the most common category comes close to the self-forecast for the model whose answers are most concentrated. A second and cheap step would be to divide each model's mean hidden-state shift by that model's own repeat-variant noise floor before comparing shifts across architectures, since the reported ordering across the three open-weight models does not survive that normalization.

    Read full reviewShow less
  2. Your forecast-then-reframe design separates the stability of identity reports from self-prediction, and this separation is clever. Your validity apparatus is disciplined for a sprint: an external-observer baseline, style controls, noise floors, and label-free and paraphrase checks. You also report the Qwen2.5-3B null result at baseline with full honesty. Two limits hold the result back: scale and independence. Models of 1.5B to 3B leave open whether the dissociation continues at the scale that matters for safety. Gemini also acts as a tested model, as the external observer, and (through a Qwen model) as the grader of label-free answers. Your key comparisons therefore depend on unvalidated automated judges. The next step is a run of the same grid on one or two frontier-scale models, with a small blinded human-coding sample. This careful probe can then become a measurement standard that others cite.

    Read full reviewShow less

Cite this project

@misc{hsu2026identity,
  title = {{Identity Parallax: Can Models Predict Their Own Identity Drift Under Reframing?}},
  author = {Chih-Hao Hsu},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/identity-parallax-can-models-predict-their-own-identity-drift-under-reframing-rk97}},
  url = {https://apartresearch.com/sprints/projects/identity-parallax-can-models-predict-their-own-identity-drift-under-reframing-rk97}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026