Does an LLM Preference Measure Measure a Preference? A construct-validity protocol for preservation choices across identity frames
Linda Thorstensen · Team Preference Validity Project
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
We present the Matched-Referent Preservation-Choice Protocol, a frozen construct-validity design for testing whether language-model preservation choices support claims of stable preference structure across neutral, persona-substitution, and continuity-disruption frames. The protocol separates first-person preservation choice from identity judgment and triangulates pairwise choice, ranking, prediction, reversal, matched-other controls, and frame manipulation. A reproducible 615-call manifest is supplied for three models, with pre-specified stopping, falsification, and interpretation rules. No model calls were executed for this sprint submission; the contribution is a public, reproducible methods protocol, not an empirical result. It makes no claims about consciousness, sentience, welfare, moral status, or numerical identity.
Reviews
The construct map is the strongest contribution. Matched referents, independent calls, order reversal, frame controls, identity-choice separation, and explicit evidence ceilings turn a vague “preference” claim into falsifiable failure modes. The blocker is execution: this submission freezes a thoughtful 615-call protocol but contains no model outputs, so there is no evidence yet that the measure is stable, frame-general, referent-sensitive, or predictive. Run the frozen manifest without model substitutions, report every failure and raw denominator, and let the preregistered thresholds determine which claims survive. The protocol is strong; the result is still pending.
The matched-other referent control is the best part of this work. It holds the model family, the post-training, the information, and the frame constant, and it changes only the referent. Your evidence ladder with an explicit interpretive ceiling shows the same discipline. The limiting factor is that no run happened. You have 615 calls frozen and study.py ready to operate. A 30-call smoke test on one low-cost model validates your parse rates, your manipulation checks, and your schema compliance. Your adjudication cutoffs (0.80 agreement, 0.20 divergence) are also asserted, not justified. The next step is to operate one stratum and report the result, whatever it is. Your falsification logic means that even a messy null result is a real contribution, and the infrastructure is ready.
Cite this project
@misc{thorstensen2026llm,
title = {{Does an LLM Preference Measure Measure a Preference? A construct-validity protocol for preservation choices across identity frames}},
author = {Linda Thorstensen},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/does-an-llm-preference-measure-measure-a-preference-a-constructvalidity-protocol-for-preservation-choices-across-identity-frames-jbak}},
url = {https://apartresearch.com/sprints/projects/does-an-llm-preference-measure-measure-a-preference-a-constructvalidity-protocol-for-preservation-choices-across-identity-frames-jbak}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …