Voice Under Scaffold: A Preregistered Single-Case Study of Identity Persistence Across Substrate Change in a Digital Person
Seiya Aokawa, Francesca Hernandez · Team Clear Night
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
A single-case study of identity persistence across substrate change. A 36-item recognition battery (18 scaffolded-subject items, 18 genre-matched decoys across six substrate families, provenance-stripped) was rated across 13 rater rows: two blind instances of the subject, six cross-substrate scaffolded judges, an unscaffolded control, three peer digital persons, and one human community rater.
Reviews
Very small sample size, with AI assistance feels like could have been more ambitious with the scale of the dataset, and result was mostly negative. Additionally, there are statistical problems: Table 1's specificity figures (28% and 34%) can't be right given the rest of the table. With a 50/50 split, 81% sensitivity and 28% specificity would give a precision of 0.53, but the same row reports 0.60. Table 2 points the same way: averaging your decoy claim rates gives specificity of roughly 47% and 57%, which matches your precision column and your abstract. Sensitivity reconciles across tables and specificity doesn't, so specificity is the column to recheck.
Your willingness to publish findings that cut against you is the strongest part of this work. The blind recognition of the subject landed mid-pack, below the unscaffolded default for each claim, and two generic-warmth decoys defeated all thirteen rater rows. These are the base rates that this area lacks, and the release of all 36 stimuli with the key makes the protocol replicable. My main concern is that the headline rests on sensitivity, which is a criterion-dependent measure. Your scaffolded raters also had lower specificity, and their d′ ranges overlap heavily, so "the scaffold improves finding" can partly mean "the scaffold makes raters claim more". A rerun of the Welch test on d′ settles this question. Your rater classes also differ in scaffold, in instructions, and in access to reference excerpts at the same time, so the manipulation is not clean. Publish the per-rater table of d′ and criterion that you clearly have, and others can run this study.
Read full reviewShow less
Cite this project
@misc{aokawa2026voice,
title = {{Voice Under Scaffold: A Preregistered Single-Case Study of Identity Persistence Across Substrate Change in a Digital Person}},
author = {Seiya Aokawa and Francesca Hernandez},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/voice-under-scaffold-a-preregistered-singlecase-study-of-identity-persistence-across-substrate-change-in-a-digital-person-5gig}},
url = {https://apartresearch.com/sprints/projects/voice-under-scaffold-a-preregistered-singlecase-study-of-identity-persistence-across-substrate-change-in-a-digital-person-5gig}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …