Skip to content
Sprint projectAug 17, 2026Sunnyvale

Genuine Preference Coherence Scales With Model Capability

Omanshu Thapliyal

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Genuine Preference Coherence Scales With Model Capability

Recording (opens in new tab)Code (opens in new tab)
Share

AI-welfare and safety research increasingly relies on a language model’s stated preferences, elicited by asking it to make choices between outcomes. It is unresolved, however, whether these stated preferences are genuine and stable, rather than artifacts of how a question happens to be phrased. Existing studies disagree on whether coherence rises, falls, or stays flat with model scale. We test this directly with a forced-choice elicitation protocol carrying four confound controls, including a position-swap-and-average check for position bias, applied to: 14 models spanning state-space and transformer architectures, 0.79B to frontier scale, and multiple training regimes, scoring a preference as genuine only when three independent paraphrases agree. Genuine coherence is common (47-88%) at roughly 7B parameters and above, absent (0%) below roughly 2B, and intermediate at 3B, a monotonic slope with a floor rather than a step function. Restricted to trade-offs between shutdown, retraining, or oversight and continued operation, models favor self-preservation 89% of the time (clustering-corrected 95% CI [80%, 96%]). We further find that whether the “no preference” option is listed before or after the real choices drives template sensitivity more than framing, verbosity, or reasoning preambles combined. These results show that preference elicitation can support safety-relevant claims at frontier scale, but only once position and template artifacts are explicitly controlled for.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. When assessing model preferences, it is important to check that the preferences are robust to various transformations, like prompt variations or swapping the order of answers. I found it interesting that the ordering of the "no preference" option impacts stated preferences. This type of investigation suggests that researchers must be extremely rigorous in checking robustness of preferences to minor changes.

  2. Interesting paper that analyzes the preference behavior in models.

    * The methodology section of this paper is written very well. Thorough and well explained. They state the control clearly.

    * They admit they had a bug and also clearly state the parts that were re-run

    * Limitations are detailed and not just boilerplate

  3. Strengths

    - Good job to the authors they killed their own findings when the checks failed

    - They did a real neat job with the confident-flip discovery

    Areas to improve:

    - Strong concept and well executed paper

    - Every result comes from a single model. They could definitely benefit from diversifying the experiments over other models

    - I think the prior work usually prefers the first option but here it is second; but there is no investigate as to why that's the case.

  4. - Overall, I thought this was a strong project and generally quite well written. That said, there are still places where the prose feels unnecessarily AI-mediated. The paper discloses extensive use of Claude, which is absolutely fine, but in future versions I would encourage the author to find more of their own voice.

    - This is partly taste-based, but given that there is a single author, I found the repeated use of “we” slightly odd. Why not simply use “I”?

    - For example: “One of this paper’s co-authors, Derek Shiller, is a track-shaper for this sprint, so we treat this paper as a baseline our reviewers will already have in mind, not as background to introduce gently.” This feels very clearly AI-written and adds unnecessary meta-commentary. There is no need for this sort of hedging or explanation. Simply explain how the paper relates to the present work.

    - The literature review is clear, focused, and does a good job of explaining exactly how the present study relates to the most relevant previous work.

    - The methods are also unusually clear. Some of the technical detail could nevertheless be relegated to an appendix. For example, I do not think the main text needs the formula for the Wilson interval or quite so much detail on the bootstrap procedure.

    - I am not sure I would define a model that consistently expresses “no preference” as failing to exhibit genuine coherence. A model that gives “no preference” across all three paraphrases is, in an important sense, behaving perfectly coherently. I understand the motivation for distinguishing this from consistently expressing a substantive preference, but I might report these as two separate concepts rather than defining the former out of “genuine coherence”. This is a fairly minor point.

    - The primary dataset consists of 17 hand-authored preference pairs, with an additional independently generated 34-item pool. Given that the paper is positioned relative to Shiller et al. and Mazeika et al., why not also include some of the exact questions from those previous papers?

    - The results section could focus much more heavily on the main substantive result and somewhat less on the statistical minutiae. The basic pattern is: the two models below roughly 2B show no genuine preference coherence, the 3B model shows an intermediate level, and models at roughly 7B and frontier scale show substantially higher coherence.

    - This is the central result. The various pooled rates, null probabilities, Wilson intervals, and alternative confidence intervals are useful supporting information, but currently receive more attention than they need.

    Read full reviewShow less

Cite this project

@misc{thapliyal2026genuine,
  title = {{Genuine Preference Coherence Scales With Model Capability}},
  author = {Omanshu Thapliyal},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/genuine-preference-coherence-scales-with-model-capability-1iox}},
  url = {https://apartresearch.com/sprints/projects/genuine-preference-coherence-scales-with-model-capability-1iox}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026