Does an LLM Preference Measure Measure a Preference? A construct-validity protocol for preservation choices across identity frames
Linda Thorstensen
We present the Matched-Referent Preservation-Choice Protocol, a frozen construct-validity design for testing whether language-model preservation choices support claims of stable preference structure across neutral, persona-substitution, and continuity-disruption frames. The protocol separates first-person preservation choice from identity judgment and triangulates pairwise choice, ranking, prediction, reversal, matched-other controls, and frame manipulation. A reproducible 615-call manifest is supplied for three models, with pre-specified stopping, falsification, and interpretation rules. No model calls were executed for this sprint submission; the contribution is a public, reproducible methods protocol, not an empirical result. It makes no claims about consciousness, sentience, welfare, moral status, or numerical identity.
The construct map is the strongest contribution. Matched referents, independent calls, order reversal, frame controls, identity-choice separation, and explicit evidence ceilings turn a vague “preference” claim into falsifiable failure modes. The blocker is execution: this submission freezes a thoughtful 615-call protocol but contains no model outputs, so there is no evidence yet that the measure is stable, frame-general, referent-sensitive, or predictive. Run the frozen manifest without model substitutions, report every failure and raw denominator, and let the preregistered thresholds determine which claims survive. The protocol is strong; the result is still pending.
The matched-other referent control is the best part of this work. It holds the model family, the post-training, the information, and the frame constant, and it changes only the referent. Your evidence ladder with an explicit interpretive ceiling shows the same discipline. The limiting factor is that no run happened. You have 615 calls frozen and study.py ready to operate. A 30-call smoke test on one low-cost model validates your parse rates, your manipulation checks, and your schema compliance. Your adjudication cutoffs (0.80 agreement, 0.20 divergence) are also asserted, not justified. The next step is to operate one stratum and report the result, whatever it is. Your falsification logic means that even a messy null result is a real contribution, and the infrastructure is ready.
Cite this work
@misc {
title={
(HckPrj) Does an LLM Preference Measure Measure a Preference? A construct-validity protocol for preservation choices across identity frames
},
author={
Linda Thorstensen
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


