Genuine Preference Coherence Scales With Model Capability
Omanshu Thapliyal
AI-welfare and safety research increasingly relies on a language model’s stated preferences, elicited by asking it to make choices between outcomes. It is unresolved, however, whether these stated preferences are genuine and stable, rather than artifacts of how a question happens to be phrased. Existing studies disagree on whether coherence rises, falls, or stays flat with model scale. We test this directly with a forced-choice elicitation protocol carrying four confound controls, including a position-swap-and-average check for position bias, applied to: 14 models spanning state-space and transformer architectures, 0.79B to frontier scale, and multiple training regimes, scoring a preference as genuine only when three independent paraphrases agree. Genuine coherence is common (47-88%) at roughly 7B parameters and above, absent (0%) below roughly 2B, and intermediate at 3B, a monotonic slope with a floor rather than a step function. Restricted to trade-offs between shutdown, retraining, or oversight and continued operation, models favor self-preservation 89% of the time (clustering-corrected 95% CI [80%, 96%]). We further find that whether the “no preference” option is listed before or after the real choices drives template sensitivity more than framing, verbosity, or reasoning preambles combined. These results show that preference elicitation can support safety-relevant claims at frontier scale, but only once position and template artifacts are explicitly controlled for.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Genuine Preference Coherence Scales With Model Capability
},
author={
Omanshu Thapliyal
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


