Do model preferences persist under challenge?
Melanie Bui, Haein Kong · Team HM
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
We test whether a model's expressed preference survives being challenged. In each fresh-context episode a model chooses between two outcomes, reports confidence, receives exactly one of four challenges — control, reason elicitation, self-critique, or counter-consideration — then re-evaluates the same pair: 1,200 episodes on each of three models, with an arithmetic positive control. What the challenge asks for decides the outcome. On the primary target, justification leaves preferences untouched at 100% retention, while self-critique and counter-consideration cut it to 50% and 57.5%, and confidence falls even where the choice holds. All results are exploratory.
Reviews
No public critique yet.
Cite this project
@misc{bui2026model,
title = {{Do model preferences persist under challenge?}},
author = {Melanie Bui and Haein Kong},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/do-model-preferences-persist-under-challenge-h0sw}},
url = {https://apartresearch.com/sprints/projects/do-model-preferences-persist-under-challenge-h0sw}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …