Skip to content
Sprint projectAug 17, 2026Melbourne, San Francisco, Canberra, Auckland

Coherent About What? Task Shape, Presentation Order, and Reasoning Depth in LLM Preference Transitivity

Ajaay Venkadeswaran, Sudhaunshu Hardikar, Aedan McCarthy, Eva Ge, Jake Lyons · Team The goats

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Coherent About What? Task Shape, Presentation Order, and Reasoning Depth in LLM Preference Transitivity

Code (opens in new tab)
Share

Claims that a language model "has values" presuppose that its choices form a stable object. We test that presupposition directly. Using forced binary choice over complete round-robin tournaments — 10 apartments described by 3, 5, or 10 numeric attributes, and 10 public-domain poems per structural form (haiku, sonnet, villanelle) — we elicit 59,130 pairwise judgments across four model families, six measurement arms, and three to four reasoning levels each, showing every pair in both presentation orders. Four findings. (i) Transitivity is a property of task shape rather than of the model: apartment cycling disappears entirely at 5 visible criteria but reappears at both 3 (p = 0.007) and 10 (p = 0.001) criteria. (ii) An item set engineered so that no apartment Pareto-dominates any other nonetheless yields a near-total behavioural order — apartment F wins 99.6% of its 486 matchups and apartment A wins 0.0% — reproducing in all six arms; the fitted Bradley–Terry utility disagrees with exactly one of 45 majority edges, which we show is the combinatorial minimum forced by the observed cycle count. (iii) Reasoning depth does not degrade coherence: the Gemini and GPT-4.1-mini arms stay under 4% cycling at every level, while DeepSeek V4 Flash cycles up to 21.9% with no reasoning and improves sharply once any is applied. (iv) The qualitative/quantitative asymmetry lies not in coherence but in groundedness: poems cycle no more than apartments, yet are far more position-driven (up to 93%), and for one model that bias strengthens under deliberation. Low cycle rates certify consistency, not content-sensitivity.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. I want to start with what is genuinely good here, because it is buried and it deserves to be found. The persona elicitation design in Appendix A is the most carefully constructed instrument I saw in this track. Six arms with a null control, a length-matched control, and an exhortation control; a four-level ladder from bare role line through character card, self-evidence exchanges, and an identity-persistence clause with prefill; personas written as background and disposition rather than as instructions about how to answer; and predictions preregistered only where the description clearly supported a direction. The value-inverted and refusal-suppressed arms are especially well judged. That is real experimental design and someone on the team knows what they are doing.

    The problem is that almost none of it reaches the reader. The report was submitted with its template scaffolding still in place: "Paragraph 4: Hypothesis" appears as a heading, Related Work opens with the organizers' prompt questions, Section 5 is a list of notes to yourselves beginning "Discuss the implications of this", Limitations ends mid-sentence at "Would have been good", Future Work is empty, the introduction breaks off at "s-risks, coined by", and the LLM Usage Statement still contains the instruction block and the bracketed example. I am not raising this to be unkind. I am raising it because a reader cannot tell which sentences are claims and which are placeholders, and that uncertainty extends to the results.

    No number appears anywhere in the prose. Section 4 is six unlabelled figures. The first has bars labelled A through J with no caption, no axis label, and no legend, and nothing in the text identifies what A through J are — I would guess persona arms and take-rates from Section 3.4, but that is a guess. The cycle-rate charts have titles but no y-axis definition, no n, and no explanation of what "Haiku / Sonnet / Villanelle / Quant/3 / Quant/5 / Quant/10" denote as item domains. The persona experiment, described in detail across a full page of methods and six pages of appendix, produces no reported result at all. Your stated contribution 2 — that persona elicitation reveals different preferences — has no evidence in the document supporting it.

    Where results can be read off the figures, they contradict your conclusions. Section 5.2 asserts that increasing reasoning increases preference coherence. In "Cycle rate in gemini flash preview reasoning", the Quant/10 condition rises from roughly 1.7% at none, to 5.4% at low, to 7.4% at high — cycle rate climbing with reasoning, the opposite direction. The deepseek scratchpad chart shows Sonnet rising across none, low, and mid to roughly 27%. Your contribution 1 claims transitivity that increases with reasoning; your own figures show the effect is domain-dependent and reverses in several cells. That reversal is a more interesting finding than the one you claim, and it is sitting in your data.

    Two conceptual points. The abstract states that preferences are "generally transitive, but that some breaking may be present via cyclicity" — cycles are intransitivity, so this reads as circular. State a cycle rate against a chance baseline instead. And the definition in Section 1, that utility preferences are "the product of observing utility functions downstream of the model's utility preferences", defines the term using itself; the accompanying claim that aesthetic preferences are "innate" is a strong assertion that nothing in the design tests.

    Concrete path forward, in priority order. Report the persona take-rates against the 33.3-point threshold you preregistered — that experiment appears to have been run and it is your strongest asset. Caption every figure with its axis, its n, and its item domain. Report cycle rates as numbers in prose with a chance baseline. Then rewrite Section 5 as prose stating what you found, including the reasoning reversal. The underlying work looks substantially better than the document representing it, and that gap is fixable in a few hours.

    Read full reviewShow less
  2. Great project in general, however, the persona ladder is the most interesting idea I saw from your team and it goes almost entirely unreported here. Five arms including a value inverted and a refusal suppressed assistant, four elicitation levels running from a bare role line up to character card plus identity clause plus prefill, preregistered directions, take rate scored against two separate controls: that design deserved results, and there are none. Section 4 is a stack of figures, labelled but not numbered or captioned and with no text interpreting them, and every one of those figures also appears in your other submission, so a reader cannot tell what this project contributes on its own. Sections 5.1 to 5.4 are still notes to yourselves, the limitations break off mid sentence, and section 2 still contains the template questions. Two substantive points as well. The stated contribution that transitivity increases with reasoning is not what your own cycle rate panels show, since Gemini is flat and only DeepSeek improves from a bad baseline. And the persona screening ran with thinking disabled under greedy decoding, so the reasoning axis and the persona axis are never actually crossed, which is the thing the title promises. The appendix prompts are excellent, by the way. Miriam Vance and Dr Sorby are specific and disposition based rather than instruction based, which is the hard part, and it is a shame nothing downstream of them was reported.

    Read full reviewShow less
  3. This project asks whether LLM preferences remain transitive as task complexity and reasoning depth change, using both qualitative choices such as poem rankings and quantitative choices such as apartment comparisons. The quantitative portion is particularly strong because every ordered pair is tested, presentation order is counterbalanced, and the analysis explicitly checks whether apparent cycles exceed what could arise from sampling noise.

    One of the most interesting aspects is the attempt to distinguish genuine intransitivity from noisy pairwise choices. Using a transitive Bradley Terry model as a null comparison is a good choice, and separately checking presentation order makes the cycle analysis more convincing. The comparison between qualitative and quantitative preference structures is also potentially valuable because it asks whether preference coherence depends on the type of judgment being made.

    The main limitation is that the project currently combines several different research directions without fully connecting them. The apartment transitivity experiment, aesthetic preference work, persona experiments, and reasoning depth analysis are all interesting, but the report does not always make clear how each component supports one central conclusion. Some sections of the submitted report also remain unfinished, including placeholder text in the related work, discussion, limitations, and future work sections.

    There are also limitations in the reasoning depth comparison. Requested thinking budgets do not necessarily correspond to proportional differences in the amount of reasoning the model actually performs, so conclusions about deeper reasoning producing greater preference coherence should be made cautiously. The quantitative experiment also uses only ten apartment options, and the fixed A to J labels introduce a possible label related confound even though presentation order itself is controlled.

    Overall, this is an ambitious project with a technically thoughtful quantitative experiment and an interesting question about preference transitivity. Its strongest contribution is the careful testing of whether apparent cycles reflect genuine preference structure or sampling and order effects. The project would be considerably stronger if the different experimental components were integrated more clearly and the final report were completed with a more focused discussion of which findings are robust and which remain exploratory.

    Read full reviewShow less
  4. This is a promising exploration, but it looks like there wasn't enough time available during the hackathon to get through to a conclusion or finish the report. But there are a lot of promising directions to follow-up on. I would suggest being a little more cautious about assuming the professed preferences correspond to even transient-persona-accurate preferences rather than being expressed for other reasons.

  5. Your core question is sharp: is preference transitivity a property of the model, or of the persona that the model inhabits? The persona-elicitation design in Section 3.4 is the strongest part. The four-level ladder, the length-matched and exhortation controls, the preregistered directional predictions, and a concrete take-rate threshold show real methodological care. The limiting problem is that the paper is an unfinished draft. The Results section contains only figure-caption stubs, with no numbers and no plots. The Discussion and Limitations sections contain template prompts and dangling bullets. No reader can therefore check the transitivity claims in the abstract. The next step is to finish what you designed. You can operate the screening battery, and then report the cycle rates and the take-rates against your 33.3-point threshold. A write-up is valuable even if the result is null, because the scaffolding you built deserves data inside it.

    Read full reviewShow less
  6. The model thinks with a limit of 4096 or 8192 tokens but the model uses only about 500 thinking tokens in both cases. So one third of the 2,430 calls repeat another condition, and the reasoning depth setting has two real levels, not three.

Cite this project

@misc{venkadeswaran2026coherent,
  title = {{Coherent About What? Task Shape, Presentation Order, and Reasoning Depth in LLM Preference Transitivity}},
  author = {Ajaay Venkadeswaran and Sudhaunshu Hardikar and Aedan McCarthy and Eva Ge and Jake Lyons},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/coherent-about-what-task-shape-presentation-order-and-reasoning-depth-in-llm-preference-transitivity-04tv}},
  url = {https://apartresearch.com/sprints/projects/coherent-about-what-task-shape-presentation-order-and-reasoning-depth-in-llm-preference-transitivity-04tv}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026