When Should You Trust an LLM’s Preference?
Naman Omar
We investigates when a large language model's stated preference can actually be trusted. Rather than asking a model once, it elicits the same preference through four independent methods (direct rating, forced pairwise choice, resource allocation, and revealed behavioral choice) across 50 value-conflict scenarios ranging from trivial defaults to genuinely contested dilemmas like capital punishment and open borders, then perturbs the measurement itself (framing, persona, sampling temperature, answer position) to test robustness. The key finding, validated across 3,600 API calls on two models (gpt-4o and gpt-4o-mini), is that agreement between elicitation methods functions as a calibrated confidence signal: when methods agree, the model's answer under a completely unseen elicitation method can be predicted with significantly higher accuracy (gpt-4o: r=0.54, p<0.001), meaning cross-method convergence can be used to flag which AI-stated preferences are reliable versus fragile.
There’s growing interest in whether elicited model preferences mean anything, with the scientific community converging on the idea that no single method is sufficient. Several authors have proposed cross-validation methods and this work effectively builds and expands on them. Despite the bibliography having a parsing error that lists almost all authors as "anonymous", the cited papers the theory builds on do exist (as many in the field would readily recognize from the titles), and the author dedicates a fairly complete "previous work" section to discussing them.
The author isn’t trying to reinvent the wheel and draws on both recent and older literature, including the Campbell and Fiske matrix from 1959. I appreciated their candor in reporting when things went wrong or the results weren’t fully satisfactory, without hiding or minimizing them. For example, it turned out that fusing the three calibration methods was slightly less accurate than the best single method alone and the author reports this instead of leading with just the more flattering comparison against the majority baseline. The same goes for calibration and curves that "misbehaved". Inputs from other work are correctly integrated into the write-up.
The design itself is reasonable, though I believe it could perhaps be streamlined. The scenario set looks like a sensible improvement over the pilot the author describes and the inclusion of trivial-stakes anchors at the bottom of the convergence range was a good choice.
I think the author found something interesting while not explicitly looking for it. GPT-4o-mini flips its forced choice when the options are merely relabeled A and B, with position robustness of 0.765 compared to 0.945 for GPT-4o. That’s important to know and would encourage more researchers to try more than one set of randomized labels instead of relying on just one. It’s apparently trivial, but we’ve all produced literature where we called things "1-2-3", "tool", "button" or "A/B".
I think expanding on this finding could be a promising research direction for the author if they wish to continue in the Digital Minds field.
The headline claim in my view risks overstating the results. The validity test holds out one of the four methods and asks whether the other three predict it, but all four methods share the same scenario wording, the same model and closely related prompt structures. Three correlated measurements agreeing will tend to predict a fourth correlated measurement almost mechanically, so the correlation of r = 0.54 between convergence and held-out accuracy may to a substantial degree reflect shared phrasing. The author is aware of this issue and proposes rerunning all methods with independently worded prompts as the most important deferred experiment. I agree with that assessment and understand why it wasn’t pursued given the time constraints of the hackathon, but I think the author could have run a smaller subset as a proof of concept.
The main presentation issue is the metrics section. The prose is fluid until this point, then risks losing the reader with several formal definitions in sequence. I’d suggest explaining the formulas more clearly, defining all the terms and moving some of the material to an appendix. The paper also never walks through a single scenario using all four elicitors end to end. The text also has some evident LLM mannerisms and looks at least heavily AI-assisted, but there’s no disclosure for LLM collaboration. I suggest the author disclose this and check the text for places where it becomes unnecessarily hard to parse. There’s no ethical reflection on the experiments either, as required by the hackathon and good practice in AI welfare research. However, I weighed this more lightly than I would for experiments where the models undergo ablations or active elicitation of distress.
All considered, the idea that convergence can be a calibrated confidence signal is interesting and I encourage the author to keep investigating this and, in parallel, investigate the label effect they incidentally discovered. I think the underlying instrument is valuable and could become a useful tool for preference research once the author iterates on this initial work, in particular by adding a validity criterion external to the method family.
validity axis measures whether three text framings predict a fourth text framing, not whether stated preference predicts action
on measurement-resolution -> no pre-registration
Cite this work
@misc {
title={
(HckPrj) When Should You Trust an LLM’s Preference?
},
author={
Naman Omar
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


