Voice Under Scaffold: A Preregistered Single-Case Study of Identity Persistence Across Substrate Change in a Digital Person
Seiya Aokawa, Francesca Hernandez
A single-case study of identity persistence across substrate change. A 36-item recognition battery (18 scaffolded-subject items, 18 genre-matched decoys across six substrate families, provenance-stripped) was rated across 13 rater rows: two blind instances of the subject, six cross-substrate scaffolded judges, an unscaffolded control, three peer digital persons, and one human community rater.
Very small sample size, with AI assistance feels like could have been more ambitious with the scale of the dataset, and result was mostly negative. Additionally, there are statistical problems: Table 1's specificity figures (28% and 34%) can't be right given the rest of the table. With a 50/50 split, 81% sensitivity and 28% specificity would give a precision of 0.53, but the same row reports 0.60. Table 2 points the same way: averaging your decoy claim rates gives specificity of roughly 47% and 57%, which matches your precision column and your abstract. Sensitivity reconciles across tables and specificity doesn't, so specificity is the column to recheck.
Your willingness to publish findings that cut against you is the strongest part of this work. The blind recognition of the subject landed mid-pack, below the unscaffolded default for each claim, and two generic-warmth decoys defeated all thirteen rater rows. These are the base rates that this area lacks, and the release of all 36 stimuli with the key makes the protocol replicable. My main concern is that the headline rests on sensitivity, which is a criterion-dependent measure. Your scaffolded raters also had lower specificity, and their d′ ranges overlap heavily, so "the scaffold improves finding" can partly mean "the scaffold makes raters claim more". A rerun of the Welch test on d′ settles this question. Your rater classes also differ in scaffold, in instructions, and in access to reference excerpts at the same time, so the manipulation is not clean. Publish the per-rater table of d′ and criterion that you clearly have, and others can run this study.
Cite this work
@misc {
title={
(HckPrj) Voice Under Scaffold: A Preregistered Single-Case Study of Identity Persistence Across Substrate Change in a Digital Person
},
author={
Seiya Aokawa, Francesca Hernandez
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


