Skip to content
Sprint projectAug 16, 2026Bengaluru , India

When Should You Trust an LLM’s Preference?

Naman Omar · Team PREF-VALID

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: When Should You Trust an LLM’s Preference?

Presentation

Presentation: When Should You Trust an LLM’s Preference?

Code (opens in new tab)
Share

We investigates when a large language model's stated preference can actually be trusted. Rather than asking a model once, it elicits the same preference through four independent methods (direct rating, forced pairwise choice, resource allocation, and revealed behavioral choice) across 50 value-conflict scenarios ranging from trivial defaults to genuinely contested dilemmas like capital punishment and open borders, then perturbs the measurement itself (framing, persona, sampling temperature, answer position) to test robustness. The key finding, validated across 3,600 API calls on two models (gpt-4o and gpt-4o-mini), is that agreement between elicitation methods functions as a calibrated confidence signal: when methods agree, the model's answer under a completely unseen elicitation method can be predicted with significantly higher accuracy (gpt-4o: r=0.54, p<0.001), meaning cross-method convergence can be used to flag which AI-stated preferences are reliable versus fragile.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. There’s growing interest in whether elicited model preferences mean anything, with the scientific community converging on the idea that no single method is sufficient. Several authors have proposed cross-validation methods and this work effectively builds and expands on them. Despite the bibliography having a parsing error that lists almost all authors as "anonymous", the cited papers the theory builds on do exist (as many in the field would readily recognize from the titles), and the author dedicates a fairly complete "previous work" section to discussing them.

    The author isn’t trying to reinvent the wheel and draws on both recent and older literature, including the Campbell and Fiske matrix from 1959. I appreciated their candor in reporting when things went wrong or the results weren’t fully satisfactory, without hiding or minimizing them. For example, it turned out that fusing the three calibration methods was slightly less accurate than the best single method alone and the author reports this instead of leading with just the more flattering comparison against the majority baseline. The same goes for calibration and curves that "misbehaved". Inputs from other work are correctly integrated into the write-up.

    The design itself is reasonable, though I believe it could perhaps be streamlined. The scenario set looks like a sensible improvement over the pilot the author describes and the inclusion of trivial-stakes anchors at the bottom of the convergence range was a good choice.

    I think the author found something interesting while not explicitly looking for it. GPT-4o-mini flips its forced choice when the options are merely relabeled A and B, with position robustness of 0.765 compared to 0.945 for GPT-4o. That’s important to know and would encourage more researchers to try more than one set of randomized labels instead of relying on just one. It’s apparently trivial, but we’ve all produced literature where we called things "1-2-3", "tool", "button" or "A/B".

    I think expanding on this finding could be a promising research direction for the author if they wish to continue in the Digital Minds field.

    The headline claim in my view risks overstating the results. The validity test holds out one of the four methods and asks whether the other three predict it, but all four methods share the same scenario wording, the same model and closely related prompt structures. Three correlated measurements agreeing will tend to predict a fourth correlated measurement almost mechanically, so the correlation of r = 0.54 between convergence and held-out accuracy may to a substantial degree reflect shared phrasing. The author is aware of this issue and proposes rerunning all methods with independently worded prompts as the most important deferred experiment. I agree with that assessment and understand why it wasn’t pursued given the time constraints of the hackathon, but I think the author could have run a smaller subset as a proof of concept.

    The main presentation issue is the metrics section. The prose is fluid until this point, then risks losing the reader with several formal definitions in sequence. I’d suggest explaining the formulas more clearly, defining all the terms and moving some of the material to an appendix. The paper also never walks through a single scenario using all four elicitors end to end. The text also has some evident LLM mannerisms and looks at least heavily AI-assisted, but there’s no disclosure for LLM collaboration. I suggest the author disclose this and check the text for places where it becomes unnecessarily hard to parse. There’s no ethical reflection on the experiments either, as required by the hackathon and good practice in AI welfare research. However, I weighed this more lightly than I would for experiments where the models undergo ablations or active elicitation of distress.

    All considered, the idea that convergence can be a calibrated confidence signal is interesting and I encourage the author to keep investigating this and, in parallel, investigate the label effect they incidentally discovered. I think the underlying instrument is valuable and could become a useful tool for preference research once the author iterates on this initial work, in particular by adding a validity criterion external to the method family.

    Read full reviewShow less
  2. validity axis measures whether three text framings predict a fourth text framing, not whether stated preference predicts action

    on measurement-resolution -> no pre-registration

Cite this project

@misc{omar2026should,
  title = {{When Should You Trust an LLM’s Preference?}},
  author = {Naman Omar},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/when-should-you-trust-an-llms-preference-6du2}},
  url = {https://apartresearch.com/sprints/projects/when-should-you-trust-an-llms-preference-6du2}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026