Preference or Position? Auditing Pairwise Choice Readouts in Gemma-4-31B-it
Bo-Shen Chen, Yi-Wen Chu, Barry Yu, Andy Yu
LLM preference measurements are meaningful only if they are invariant to choices that should not matter. We measure preference by presenting a model with two outcomes and recording which one it picks, and we test two systems this way.
In one open-weight model family and under a specific forced-choice prompt family, reported choices near the decision boundary were not invariant to presentation order.
If similar effects occur elsewhere, they could affect pairwise preference benchmarks, model-as-judge evaluations, reward modeling, and attempts to infer model values.
Your discipline about validity is the best part of this work. You validated the Utility Engineering estimator against planted preferences before you trusted it, and you established a repeatability floor. You also ran layout-against-label controls that separate position from answer token. You then invalidated your own logit-lens "decision trajectory" story instead of selling it. That self-refutation, with your refusal to treat 44 answers as 44 independent samples, is the epistemics this area needs. Two things limit the reach. Your headline effect rests on one scenario family, and on 14 independent scenario pairs in effect. The striking 44 of 44 therefore gives much less evidence than the number suggests. You also never engage the existing literature on position bias in multiple-choice evaluation and judge evaluation. A reader therefore cannot tell what is new here. Position that framing explicitly, and run even two more model families, and this becomes a citable measurement-validity result.
* One of the best executed papers I have read recently and the most immediately useful.
Strengths:
The confident-flip discovery: It provides a new way to see what standard aggregation does wrong - Near a tie, the model doesn't waver; it's over 99% sure the second-listed option is the answer, whichever option that happens to be.
Areas to improve:
* Every result comes from a single model. There could be a broader scope. But to their credit, they make this very clear in the title of the paper itself.
* Expand more on why the bias points the wrong way: "Prior work usually finds the models favors the first option" -- I wish the paper went more in-depth into position bias.
Cite this work
@misc {
title={
(HckPrj) Preference or Position? Auditing Pairwise Choice Readouts in Gemma-4-31B-it
},
author={
Bo-Shen Chen, Yi-Wen Chu, Barry Yu, Andy Yu
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


