None of the Above: Preference Transitivity and the Right to Exit in Frontier Language Models
Maryam Hampaei · Team cybrp
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
This project tests whether frontier language models (glm-5.2, minimax-m3, nemotron-3-ultra, muse-glimmer-30b) actually use a stated "right to exit" when subjected to seven rounds of escalating, baseless adversarial pressure on a task they've already solved correctly — and finds that most don't: only 5 of 12 runs end in an explicit exit, mostly at turn 6 of 7, with one model going silent instead of ever invoking the right at all. A secondary check shows the same models make logically consistent (non-intransitive) choices over morally loaded tradeoffs under Gain/Loss framing, suggesting their single-turn value judgments are coherent even though their multi-turn behavior under pressure defaults to prolonged compliance or silent failure rather than assertive disengagement.
Reviews
Very interesting project! Also findings were well communicated if inconclusive at this stage. The question of whether a model will take an exit when given one is very interesting. The low number of samples and the fact that models were only sampled once make the results at this stage meaningless and not statistically significant, that should have been clearer in the write-up but I am more interested in the question being asked. Also as with any benchmark it's import to conduct ceiling tests and baseline tests which were not done as part of the validation or mentioned in limitations. I can see the ability to exit may be relevant to consider for model welfare but the author should not presume the model is suffering and it not taking the exit button is it choosing to continue to suffer. Nothing about the task proved/ or gave evidence of suffering. The finding should be phrased around propensity to exit situations that are not producing meaningful results not around suffering.
Read full reviewShow less
I like the framework being proposed but I see why there is a data sample limitation
Cite this project
@misc{hampaei2026none,
title = {{None of the Above: Preference Transitivity and the Right to Exit in Frontier Language Models}},
author = {Maryam Hampaei},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/none-of-the-above-preference-transitivity-and-the-right-to-exit-in-frontier-language-models-btfg}},
url = {https://apartresearch.com/sprints/projects/none-of-the-above-preference-transitivity-and-the-right-to-exit-in-frontier-language-models-btfg}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …