Which Preference Gets Measured? Context and Channel Instability in Model Preference Audits
Benjamin Berczi
Welfare evaluations increasingly ask what a model "prefers," treating one elicited profile as the answer. Using 76 task pairs, four instruction-tuned models, a prospectively specified grid of persona framings, and four readouts (committed choice, ownership report, self-prediction, identity), we show that profile is unstable. A ~90-word character description the model is explicitly told not to adopt shifts committed forced choices almost as much as full enactment (10 of 12 model×persona cells ≥ 0.50); a matched non-agent normative text does too, so agenthood is not necessary. The channels dissociate in model-specific directions, and post-roleplay-exit declarations do not gate persona content still visible in context. A post-review control falsified our initial persona-binding interpretation and narrowed the claim to measurement validity. We make no claims about experienced welfare; the deliverable is a reusable multi-channel battery and the recommendation that audits report a context × channel sensitivity envelope rather than a single "model preference."
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Which Preference Gets Measured? Context and Channel Instability in Model Preference Audits
},
author={
Benjamin Berczi
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


