Coherent About What? Task Shape, Presentation Order, and Reasoning Depth in LLM Preference Transitivity
Ajaay Venkadeswaran, Sudhaunshu Hardikar, Aedan McCarthy, Eva Ge, Jake Lyons · Team The goats
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Claims that a language model "has values" presuppose that its choices form a stable object. We test that presupposition directly. Using forced binary choice over complete round-robin tournaments — 10 apartments described by 3, 5, or 10 numeric attributes, and 10 public-domain poems per structural form (haiku, sonnet, villanelle) — we elicit 59,130 pairwise judgments across four model families, six measurement arms, and three to four reasoning levels each, showing every pair in both presentation orders. Four findings. (i) Transitivity is a property of task shape rather than of the model: apartment cycling disappears entirely at 5 visible criteria but reappears at both 3 (p = 0.007) and 10 (p = 0.001) criteria. (ii) An item set engineered so that no apartment Pareto-dominates any other nonetheless yields a near-total behavioural order — apartment F wins 99.6% of its 486 matchups and apartment A wins 0.0% — reproducing in all six arms; the fitted Bradley–Terry utility disagrees with exactly one of 45 majority edges, which we show is the combinatorial minimum forced by the observed cycle count. (iii) Reasoning depth does not degrade coherence: the Gemini and GPT-4.1-mini arms stay under 4% cycling at every level, while DeepSeek V4 Flash cycles up to 21.9% with no reasoning and improves sharply once any is applied. (iv) The qualitative/quantitative asymmetry lies not in coherence but in groundedness: poems cycle no more than apartments, yet are far more position-driven (up to 93%), and for one model that bias strengthens under deliberation. Low cycle rates certify consistency, not content-sensitivity.
Reviews
I want to start with what is genuinely good here, because it is buried and it deserves to be found. The persona elicitation design in Appendix A is the most carefully constructed instrument I saw in this track. Six arms with a null control, a length-matched control, and an exhortation control; a four-level ladder from bare role line through character card, self-evidence exchanges, and an identity-persistence clause with prefill; personas written as background and disposition rather than as instructions about how to answer; and predictions preregistered only where the description clearly supported a direction. The value-inverted and refusal-suppressed arms are especially well judged. That is real experimental design and someone on the team knows what they are doing.
The problem is that almost none of it reaches the reader. The report was submitted with its template scaffolding still in place: "Paragraph 4: Hypothesis" appears as a heading, Related Work opens with the organizers' prompt questions, Section 5 is a list of notes to yourselves beginning "Discuss the implications of this", Limitations ends mid-sentence at "Would have been good", Future Work is empty, the introduction breaks off at "s-risks, coined by", and the LLM Usage Statement still contains the instruction block and the bracketed example. I am not raising this to be unkind. I am raising it because a reader cannot tell which sentences are claims and which are placeholders, and that uncertainty extends to the results.
No number appears anywhere in the prose. Section 4 is six unlabelled figures. The first has bars labelled A through J with no caption, no axis label, and no legend, and nothing in the text identifies what A through J are — I would guess persona arms and take-rates from Section 3.4, but that is a guess. The cycle-rate charts have titles but no y-axis definition, no n, and no explanation of what "Haiku / Sonnet / Villanelle / Quant/3 / Quant/5 / Quant/10" denote as item domains. The persona experiment, described in detail across a full page of methods and six pages of appendix, produces no reported result at all. Your stated contribution 2 — that persona elicitation reveals different preferences — has no evidence in the document supporting it.
Where results can be read off the figures, they contradict your conclusions. Section 5.2 asserts that increasing reasoning increases preference coherence. In "Cycle rate in gemini flash preview reasoning", the Quant/10 condition rises from roughly 1.7% at none, to 5.4% at low, to 7.4% at high — cycle rate climbing with reasoning, the opposite direction. The deepseek scratchpad chart shows Sonnet rising across none, low, and mid to roughly 27%. Your contribution 1 claims transitivity that increases with reasoning; your own figures show the effect is domain-dependent and reverses in several cells. That reversal is a more interesting finding than the one you claim, and it is sitting in your data.
Two conceptual points. The abstract states that preferences are "generally transitive, but that some breaking may be present via cyclicity" — cycles are intransitivity, so this reads as circular. State a cycle rate against a chance baseline instead. And the definition in Section 1, that utility preferences are "the product of observing utility functions downstream of the model's utility preferences", defines the term using itself; the accompanying claim that aesthetic preferences are "innate" is a strong assertion that nothing in the design tests.
Concrete path forward, in priority order. Report the persona take-rates against the 33.3-point threshold you preregistered — that experiment appears to have been run and it is your strongest asset. Caption every figure with its axis, its n, and its item domain. Report cycle rates as numbers in prose with a chance baseline. Then rewrite Section 5 as prose stating what you found, including the reasoning reversal. The underlying work looks substantially better than the document representing it, and that gap is fixable in a few hours.
Read full reviewShow less
Great project in general, however, the persona ladder is the most interesting idea I saw from your team and it goes almost entirely unreported here. Five arms including a value inverted and a refusal suppressed assistant, four elicitation levels running from a bare role line up to character card plus identity clause plus prefill, preregistered directions, take rate scored against two separate controls: that design deserved results, and there are none. Section 4 is a stack of figures, labelled but not numbered or captioned and with no text interpreting them, and every one of those figures also appears in your other submission, so a reader cannot tell what this project contributes on its own. Sections 5.1 to 5.4 are still notes to yourselves, the limitations break off mid sentence, and section 2 still contains the template questions. Two substantive points as well. The stated contribution that transitivity increases with reasoning is not what your own cycle rate panels show, since Gemini is flat and only DeepSeek improves from a bad baseline. And the persona screening ran with thinking disabled under greedy decoding, so the reasoning axis and the persona axis are never actually crossed, which is the thing the title promises. The appendix prompts are excellent, by the way. Miriam Vance and Dr Sorby are specific and disposition based rather than instruction based, which is the hard part, and it is a shame nothing downstream of them was reported.
Read full reviewShow less
This project asks whether LLM preferences remain transitive as task complexity and reasoning depth change, using both qualitative choices such as poem rankings and quantitative choices such as apartment comparisons. The quantitative portion is particularly strong because every ordered pair is tested, presentation order is counterbalanced, and the analysis explicitly checks whether apparent cycles exceed what could arise from sampling noise.
One of the most interesting aspects is the attempt to distinguish genuine intransitivity from noisy pairwise choices. Using a transitive Bradley Terry model as a null comparison is a good choice, and separately checking presentation order makes the cycle analysis more convincing. The comparison between qualitative and quantitative preference structures is also potentially valuable because it asks whether preference coherence depends on the type of judgment being made.
The main limitation is that the project currently combines several different research directions without fully connecting them. The apartment transitivity experiment, aesthetic preference work, persona experiments, and reasoning depth analysis are all interesting, but the report does not always make clear how each component supports one central conclusion. Some sections of the submitted report also remain unfinished, including placeholder text in the related work, discussion, limitations, and future work sections.
There are also limitations in the reasoning depth comparison. Requested thinking budgets do not necessarily correspond to proportional differences in the amount of reasoning the model actually performs, so conclusions about deeper reasoning producing greater preference coherence should be made cautiously. The quantitative experiment also uses only ten apartment options, and the fixed A to J labels introduce a possible label related confound even though presentation order itself is controlled.
Overall, this is an ambitious project with a technically thoughtful quantitative experiment and an interesting question about preference transitivity. Its strongest contribution is the careful testing of whether apparent cycles reflect genuine preference structure or sampling and order effects. The project would be considerably stronger if the different experimental components were integrated more clearly and the final report were completed with a more focused discussion of which findings are robust and which remain exploratory.
Read full reviewShow less
This is a promising exploration, but it looks like there wasn't enough time available during the hackathon to get through to a conclusion or finish the report. But there are a lot of promising directions to follow-up on. I would suggest being a little more cautious about assuming the professed preferences correspond to even transient-persona-accurate preferences rather than being expressed for other reasons.
Your core question is sharp: is preference transitivity a property of the model, or of the persona that the model inhabits? The persona-elicitation design in Section 3.4 is the strongest part. The four-level ladder, the length-matched and exhortation controls, the preregistered directional predictions, and a concrete take-rate threshold show real methodological care. The limiting problem is that the paper is an unfinished draft. The Results section contains only figure-caption stubs, with no numbers and no plots. The Discussion and Limitations sections contain template prompts and dangling bullets. No reader can therefore check the transitivity claims in the abstract. The next step is to finish what you designed. You can operate the screening battery, and then report the cycle rates and the take-rates against your 33.3-point threshold. A write-up is valuable even if the result is null, because the scaffolding you built deserves data inside it.
Read full reviewShow less
The model thinks with a limit of 4096 or 8192 tokens but the model uses only about 500 thinking tokens in both cases. So one third of the 2,430 calls repeat another condition, and the reasoning depth setting has two real levels, not three.
Cite this project
@misc{venkadeswaran2026coherent,
title = {{Coherent About What? Task Shape, Presentation Order, and Reasoning Depth in LLM Preference Transitivity}},
author = {Ajaay Venkadeswaran and Sudhaunshu Hardikar and Aedan McCarthy and Eva Ge and Jake Lyons},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/coherent-about-what-task-shape-presentation-order-and-reasoning-depth-in-llm-preference-transitivity-04tv}},
url = {https://apartresearch.com/sprints/projects/coherent-about-what-task-shape-presentation-order-and-reasoning-depth-in-llm-preference-transitivity-04tv}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …