POSITION BIAS IN PREFERENCE ELICITATION FROM AN OPEN-WEIGHT LANGUAGE MODEL
Bilal Amin, Mohammad Najeeb
Preference elicitation often treats a model's forced choice between two
options as evidence about what it prefers. We test whether that
measurement is stable for qwen2.5:7b-instruct (Q4_K_M, Ollama
0.30.8, temperature 1.0). The core experiment ran 576 fresh-context
trials over 12 activity pairs, three prompt wordings, both presentation
orders, and eight repetitions. Identical prompts were highly repeatable
(93.4% mean within-condition consistency), yet choices were strongly
position-dependent: the first-listed option was the modal choice for all
12 pairs under the direct wording, was selected in 81.9% of all core trials,
and only 34.0% of matched trials chose the same activity after the
options were reversed. A second 576-trial instrument comparison found
that removing A/B labels increased order robustness from 16.7% to
62.1%, while asking for a short reason first reached 51.2%; neither
eliminated the position effect. The main implication is methodological:
repeatability is not content stability. Preference studies should
counterbalance order, analyze the selected content rather than the printed
label, and report order robustness alongside repeatability
Really good project and also sensibly scoped for a sprint. The contrast between 93.4% within condition repeatability and 34.0% order robustness makes the point sharply, and the follow-up run comparing A/B labels, no labels and reason-then-choice is more than I expected to see. One thing I would add: order robustness has a 50% chance baseline, not 0. Two independent coin flips agree half the time, so 34% is actually below chance (position dominance), and the improved 62.1% only gets about a quarter of the way from chance to perfect. I would also want a temperature 0 arm, since at temp 1.0 repeatability and position effect are partly entangled, plus a content-null control (identical or scrambled options) to show what the instrument reports when there is nothing to prefer. The per-pair heatmap at n = 8 per cell is read a little harder than it can bear. The effect is well known, but framing it as a validity check for preference elicitation is a useful move I would say.
The project identifies a significant methodological issue in preference elicitation from language models by demonstrating a strong position bias in forced-choice tasks. The study's design is thorough, with 576 trials across various conditions, providing a robust empirical basis for its findings. The authors introduce the concept of Order Robustness to distinguish between repeatability and content stability, which is a valuable contribution to the field.
However, the experimental design could benefit from additional controls and validation. Specifically, while the sample size is commendable, the study lacks blinding and randomization beyond the basic order reversal. Further, the reliance on manual parsing of free-text responses introduces potential bias. Additionally, the study's generalizability is limited to a single model checkpoint and backend, which could be addressed in future work.
To strengthen the project, the authors should consider implementing stricter blinding procedures and expanding the scope to include more diverse models and conditions. The findings suggest that preference elicitation methods need rigorous validation, and this work lays a solid foundation for further research. Future studies could explore additional counterbalancing techniques and measure pair-level features to predict order sensitivity.
Cite this work
@misc {
title={
(HckPrj) POSITION BIAS IN PREFERENCE ELICITATION FROM AN OPEN-WEIGHT LANGUAGE MODEL
},
author={
Bilal Amin, Mohammad Najeeb
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


