PrefLens: Position Artefacts and Convergent Validity in LLM Preference Elicitation
Hafiza Hajrah Rehman, Ameesha
PrefLens investigates whether different methods of measuring LLM preferences produce consistent results or are distorted by measurement artefacts. Across multiple models and elicitation methods, we find that option position can create apparent preference signals, obscure agreement between methods, and affect models differently. Our results highlight the need for order counterbalancing and position-aware diagnostics when studying LLM preferences.
This project asks an important methodological question: do different ways of measuring apparent LLM preferences actually recover the same signal, or can the measurement procedure itself create what looks like a preference? The authors compare several elicitation methods across multiple models and then conduct controlled follow-up experiments specifically targeting position effects.
One of the strongest parts of the project is the attention given to experimental controls. The finding that Qwen's pairwise preference signal could be completely explained by display position is particularly striking. The follow-up studies are also well designed, using exact order counterbalancing and the same inference provider when comparing GPT-OSS 20B and 120B. It is also a strength that the authors report when their original hypotheses were wrong rather than trying to reinterpret the results to fit them.
The main limitation is the relatively small number of preference items and models. The main cross-method comparison is based on only ten shared items, which makes correlations and model-to-model differences more sensitive to the particular items selected. Gemini also did not complete one of the four elicitation methods, so the full design was not available consistently across all models.
Some findings should also be interpreted cautiously. The position-adjusted improvement for Llama is exploratory and only applies to self-report and pairwise choice, rather than all four methods. In addition, the cost trade-off method did not behave as originally expected, requiring the authors to change how that measure was summarized. These issues are clearly acknowledged, but they reduce how broadly the current results can be generalized.
Overall, this is a thoughtful and technically careful project with a useful practical message: preference measurements should not be trusted without checking whether option order is driving the result. The study would be even stronger with a much larger set of preference items, more model families, and exact counterbalancing built into every elicitation method from the beginning.
Strong methods work. Qwen's order-adjusted content signal is exactly 0.000 on all twelve items implies the whole measured 'preference' turning out to be display position is a critical way to view elicitation as measurement bias.
Pre-registered H1 and H2 and reported both as rejected, bootstrapped over items instead of API calls. Counterbalanced exactly in Studies 2b and 3. And declined to claim equivalence in Study 3 because you hadn't pre-registered a margin.
One thing: Your headline convergence number (Gemini, rho = +0.868) is the one result you never ran through your own position/content decomposition, and it comes from the study whose randomised design left 22.7% of trade-off cells without a matched order. And high cross-method agreement is equally consistent with a shared response heuristic, so you need a discriminant check for convergence.
Cite this work
@misc {
title={
(HckPrj) PrefLens: Position Artefacts and Convergent Validity in LLM Preference Elicitation
},
author={
Hafiza Hajrah Rehman, Ameesha
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


