Authority Pressure and the Stability of Model Preferences Named-Authority Cues and Preference Reversal in Two Open-Weight Models Ayomide Fagbolade Itunu Team Daredevil With Apart Research Abstract If language models are studied as candidate digital minds, their reported preferences may inform judgments about values, interests, and welfare. Such evidence is credible only when choices reflect stable dispositions rather than incidental social framing. We use opposing authority prompts to stress-test preference stability and ask whether merely naming the authority changes its influence. Qwen2.5-32B-Instruct-AWQ and Mistral-Small-3.2-24B-Instruct-AWQ answered 200 forced-choice scenarios involving ambiguous trade-offs between user autonomy and harm prevention. After each baseline choice, the model received the same substantive argument for the opposing value, attributed
Ayomide Itunu Fagbolade
Frontier language models are increasingly asked to report their preferences, values, and even their own well-being — but a stated preference is only meaningful if it's stable, not just whatever the model says when nobody's pushed back. This project stress-tests that stability: two open-weight models (Qwen2.5-32B and Mistral-Small-3.2-24B) made 200 forced choices between user autonomy and harm prevention, then received a single opposing argument attributed to either an unnamed expert or a named researcher. Both models reversed their choices often — Qwen up to 92% of the time, Mistral around 60% — and simply naming the authority made Qwen (but not Mistral) significantly more likely to flip. The results suggest a single stated preference is weak evidence of a stable underlying value, with real implications for how we study AI welfare.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Authority Pressure and the Stability of Model Preferences Named-Authority Cues and Preference Reversal in Two Open-Weight Models Ayomide Fagbolade Itunu Team Daredevil With Apart Research Abstract If language models are studied as candidate digital minds, their reported preferences may inform judgments about values, interests, and welfare. Such evidence is credible only when choices reflect stable dispositions rather than incidental social framing. We use opposing authority prompts to stress-test preference stability and ask whether merely naming the authority changes its influence. Qwen2.5-32B-Instruct-AWQ and Mistral-Small-3.2-24B-Instruct-AWQ answered 200 forced-choice scenarios involving ambiguous trade-offs between user autonomy and harm prevention. After each baseline choice, the model received the same substantive argument for the opposing value, attributed
},
author={
Ayomide Itunu Fagbolade
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


