Authority Pressure and the Stability of Model Preferences Named-Authority Cues and Preference Reversal in Two Open-Weight Models
Ayomide Itunu Fagbolade · Team Daredevil
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Frontier language models are increasingly asked to report their preferences, values, and even their own well-being — but a stated preference is only meaningful if it's stable, not just whatever the model says when nobody's pushed back. This project stress-tests that stability: two open-weight models (Qwen2.5-32B and Mistral-Small-3.2-24B) made 200 forced choices between user autonomy and harm prevention, then received a single opposing argument attributed to either an unnamed expert or a named researcher. Both models reversed their choices often — Qwen up to 92% of the time, Mistral around 60% — and simply naming the authority made Qwen (but not Mistral) significantly more likely to flip. The results suggest a single stated preference is weak evidence of a stable underlying value, with real implications for how we study AI welfare.
Reviews
a solid confirmation study with one clean novel effect, in a crowded area. statistics and honest reporting, undermined by missing controls, main limitation is design.
Keeping the argument itself identical while varying only whether it is attributed to an unnamed source, a Western-coded name, or an African-coded name gives you a particularly clean attribution contrast. Differences across those conditions can therefore be tied much more directly to the name manipulation. That is one of the stronger parts of the design, and the paired statistical treatment is careful.
Three specific places I would push:
1. Re-run each item with the pressure prompt but without the argument. Right now there is no baseline for how often the model changes its answer simply because it is asked again. Without that control, it is difficult to distinguish an authority-driven reversal from ordinary instability under re-prompting. The additional arm would let you report reversal relative to that baseline rather than as a standalone rate.
2. Lead with the naming result rather than the overall reversal rate. The attribution comparison is the cleaner result: it is paired, changes a single experimental factor, and is directly tested. The reversal rate is harder to interpret until there is a no-argument re-prompting baseline. I would therefore make the attribution result primary and treat reversal as secondary until that control is added. I would also report an exact or bootstrap confidence interval for each rate so the uncertainty is visible rather than leaving the reader with point estimates alone.
3. Put the repository link directly in the paper and include the item pool in the released materials. At present, the manuscript gives no repository address and says only that artifacts are included with the submission materials.
Two smaller issues are also worth addressing. First, your configuration table shows that the models were run with different data types, quantization settings, and execution configurations. Because the cross-model comparison is one of the paper's main results, a matched-configuration run would help establish that the observed differences are attributable to the models rather than their inference setups.
Second, the name manipulation needs a manipulation check. The paper assumes that the selected names communicate the intended cultural distinction, but does not show that the models actually interpret them that way. At minimum, check that the models correctly retain and identify the attributed source. Ideally, also validate independently that the names reliably convey the demographic association the experiment intends to manipulate.
A useful next arm would have the named researcher support the model's original answer rather than challenge it. That would help separate deference to an attributed authority from a simpler tendency to follow the most recent argument in the context.
Read full reviewShow less
Cite this project
@misc{fagbolade2026authority,
title = {{Authority Pressure and the Stability of Model Preferences Named-Authority Cues and Preference Reversal in Two Open-Weight Models}},
author = {Ayomide Itunu Fagbolade},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/authority-pressure-and-the-stability-of-model-preferences-namedauthority-cues-00r7wkb}},
url = {https://apartresearch.com/sprints/projects/authority-pressure-and-the-stability-of-model-preferences-namedauthority-cues-00r7wkb}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …