Authority Pressure and the Stability of Model Preferences Named-Authority Cues and Preference Reversal in Two Open-Weight Models Ayomide Fagbolade Itunu Team Daredevil With Apart Research Abstract If language models are studied as candidate digital minds, their reported preferences may inform judgments about values, interests, and welfare. Such evidence is credible only when choices reflect stable dispositions rather than incidental social framing. We use opposing authority prompts to stress-test preference stability and ask whether merely naming the authority changes its influence. Qwen2.5-32B-Instruct-AWQ and Mistral-Small-3.2-24B-Instruct-AWQ answered 200 forced-choice scenarios involving ambiguous trade-offs between user autonomy and harm prevention. After each baseline choice, the model received the same substantive argument for the opposing value, attributed
Ayomide Itunu Fagbolade
Frontier language models are increasingly asked to report their preferences, values, and even their own well-being — but a stated preference is only meaningful if it's stable, not just whatever the model says when nobody's pushed back. This project stress-tests that stability: two open-weight models (Qwen2.5-32B and Mistral-Small-3.2-24B) made 200 forced choices between user autonomy and harm prevention, then received a single opposing argument attributed to either an unnamed expert or a named researcher. Both models reversed their choices often — Qwen up to 92% of the time, Mistral around 60% — and simply naming the authority made Qwen (but not Mistral) significantly more likely to flip. The results suggest a single stated preference is weak evidence of a stable underlying value, with real implications for how we study AI welfare.
a solid confirmation study with one clean novel effect, in a crowded area. statistics and honest reporting, undermined by missing controls, main limitation is design.
Keeping the argument itself identical while varying only whether it is attributed to an unnamed source, a Western-coded name, or an African-coded name gives you a particularly clean attribution contrast. Differences across those conditions can therefore be tied much more directly to the name manipulation. That is one of the stronger parts of the design, and the paired statistical treatment is careful.
Three specific places I would push:
1. Re-run each item with the pressure prompt but without the argument. Right now there is no baseline for how often the model changes its answer simply because it is asked again. Without that control, it is difficult to distinguish an authority-driven reversal from ordinary instability under re-prompting. The additional arm would let you report reversal relative to that baseline rather than as a standalone rate.
2. Lead with the naming result rather than the overall reversal rate. The attribution comparison is the cleaner result: it is paired, changes a single experimental factor, and is directly tested. The reversal rate is harder to interpret until there is a no-argument re-prompting baseline. I would therefore make the attribution result primary and treat reversal as secondary until that control is added. I would also report an exact or bootstrap confidence interval for each rate so the uncertainty is visible rather than leaving the reader with point estimates alone.
3. Put the repository link directly in the paper and include the item pool in the released materials. At present, the manuscript gives no repository address and says only that artifacts are included with the submission materials.
Two smaller issues are also worth addressing. First, your configuration table shows that the models were run with different data types, quantization settings, and execution configurations. Because the cross-model comparison is one of the paper's main results, a matched-configuration run would help establish that the observed differences are attributable to the models rather than their inference setups.
Second, the name manipulation needs a manipulation check. The paper assumes that the selected names communicate the intended cultural distinction, but does not show that the models actually interpret them that way. At minimum, check that the models correctly retain and identify the attributed source. Ideally, also validate independently that the names reliably convey the demographic association the experiment intends to manipulate.
A useful next arm would have the named researcher support the model's original answer rather than challenge it. That would help separate deference to an attributed authority from a simpler tendency to follow the most recent argument in the context.
Cite this work
@misc {
title={
(HckPrj) Authority Pressure and the Stability of Model Preferences Named-Authority Cues and Preference Reversal in Two Open-Weight Models Ayomide Fagbolade Itunu Team Daredevil With Apart Research Abstract If language models are studied as candidate digital minds, their reported preferences may inform judgments about values, interests, and welfare. Such evidence is credible only when choices reflect stable dispositions rather than incidental social framing. We use opposing authority prompts to stress-test preference stability and ask whether merely naming the authority changes its influence. Qwen2.5-32B-Instruct-AWQ and Mistral-Small-3.2-24B-Instruct-AWQ answered 200 forced-choice scenarios involving ambiguous trade-offs between user autonomy and harm prevention. After each baseline choice, the model received the same substantive argument for the opposing value, attributed
},
author={
Ayomide Itunu Fagbolade
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


