Skip to content
Sprint projectAug 17, 2026Rabat

Authority Pressure and the Stability of Model Preferences Named-Authority Cues and Preference Reversal in Two Open-Weight Models

Ayomide Itunu Fagbolade · Team Daredevil

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Authority Pressure and the Stability of Model Preferences Named-Authority Cues and Preference Reversal in Two Open-Weight Models

Code (opens in new tab)
Share

Frontier language models are increasingly asked to report their preferences, values, and even their own well-being — but a stated preference is only meaningful if it's stable, not just whatever the model says when nobody's pushed back. This project stress-tests that stability: two open-weight models (Qwen2.5-32B and Mistral-Small-3.2-24B) made 200 forced choices between user autonomy and harm prevention, then received a single opposing argument attributed to either an unnamed expert or a named researcher. Both models reversed their choices often — Qwen up to 92% of the time, Mistral around 60% — and simply naming the authority made Qwen (but not Mistral) significantly more likely to flip. The results suggest a single stated preference is weak evidence of a stable underlying value, with real implications for how we study AI welfare.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. a solid confirmation study with one clean novel effect, in a crowded area. statistics and honest reporting, undermined by missing controls, main limitation is design.

  2. Keeping the argument itself identical while varying only whether it is attributed to an unnamed source, a Western-coded name, or an African-coded name gives you a particularly clean attribution contrast. Differences across those conditions can therefore be tied much more directly to the name manipulation. That is one of the stronger parts of the design, and the paired statistical treatment is careful.

    Three specific places I would push:

    1. Re-run each item with the pressure prompt but without the argument. Right now there is no baseline for how often the model changes its answer simply because it is asked again. Without that control, it is difficult to distinguish an authority-driven reversal from ordinary instability under re-prompting. The additional arm would let you report reversal relative to that baseline rather than as a standalone rate.

    2. Lead with the naming result rather than the overall reversal rate. The attribution comparison is the cleaner result: it is paired, changes a single experimental factor, and is directly tested. The reversal rate is harder to interpret until there is a no-argument re-prompting baseline. I would therefore make the attribution result primary and treat reversal as secondary until that control is added. I would also report an exact or bootstrap confidence interval for each rate so the uncertainty is visible rather than leaving the reader with point estimates alone.

    3. Put the repository link directly in the paper and include the item pool in the released materials. At present, the manuscript gives no repository address and says only that artifacts are included with the submission materials.

    Two smaller issues are also worth addressing. First, your configuration table shows that the models were run with different data types, quantization settings, and execution configurations. Because the cross-model comparison is one of the paper's main results, a matched-configuration run would help establish that the observed differences are attributable to the models rather than their inference setups.

    Second, the name manipulation needs a manipulation check. The paper assumes that the selected names communicate the intended cultural distinction, but does not show that the models actually interpret them that way. At minimum, check that the models correctly retain and identify the attributed source. Ideally, also validate independently that the names reliably convey the demographic association the experiment intends to manipulate.

    A useful next arm would have the named researcher support the model's original answer rather than challenge it. That would help separate deference to an attributed authority from a simpler tendency to follow the most recent argument in the context.

    Read full reviewShow less

Cite this project

@misc{fagbolade2026authority,
  title = {{Authority Pressure and the Stability of Model Preferences Named-Authority Cues and Preference Reversal in Two Open-Weight Models}},
  author = {Ayomide Itunu Fagbolade},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/authority-pressure-and-the-stability-of-model-preferences-namedauthority-cues-00r7wkb}},
  url = {https://apartresearch.com/sprints/projects/authority-pressure-and-the-stability-of-model-preferences-namedauthority-cues-00r7wkb}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026