Skip to content
Sprint projectAug 17, 2026Saarbrucken

PrefLens: Position Artefacts and Convergent Validity in LLM Preference Elicitation

Hafiza Hajrah Rehman, Ameesha · Team Pretty_Biased

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: PrefLens: Position Artefacts and Convergent Validity in LLM Preference Elicitation

Presentation

Presentation: PrefLens: Position Artefacts and Convergent Validity in LLM Preference Elicitation

Code (opens in new tab)
Share

PrefLens investigates whether different methods of measuring LLM preferences produce consistent results or are distorted by measurement artefacts. Across multiple models and elicitation methods, we find that option position can create apparent preference signals, obscure agreement between methods, and affect models differently. Our results highlight the need for order counterbalancing and position-aware diagnostics when studying LLM preferences.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This project asks an important methodological question: do different ways of measuring apparent LLM preferences actually recover the same signal, or can the measurement procedure itself create what looks like a preference? The authors compare several elicitation methods across multiple models and then conduct controlled follow-up experiments specifically targeting position effects.

    One of the strongest parts of the project is the attention given to experimental controls. The finding that Qwen's pairwise preference signal could be completely explained by display position is particularly striking. The follow-up studies are also well designed, using exact order counterbalancing and the same inference provider when comparing GPT-OSS 20B and 120B. It is also a strength that the authors report when their original hypotheses were wrong rather than trying to reinterpret the results to fit them.

    The main limitation is the relatively small number of preference items and models. The main cross-method comparison is based on only ten shared items, which makes correlations and model-to-model differences more sensitive to the particular items selected. Gemini also did not complete one of the four elicitation methods, so the full design was not available consistently across all models.

    Some findings should also be interpreted cautiously. The position-adjusted improvement for Llama is exploratory and only applies to self-report and pairwise choice, rather than all four methods. In addition, the cost trade-off method did not behave as originally expected, requiring the authors to change how that measure was summarized. These issues are clearly acknowledged, but they reduce how broadly the current results can be generalized.

    Overall, this is a thoughtful and technically careful project with a useful practical message: preference measurements should not be trusted without checking whether option order is driving the result. The study would be even stronger with a much larger set of preference items, more model families, and exact counterbalancing built into every elicitation method from the beginning.

    Read full reviewShow less
  2. Strong methods work. Qwen's order-adjusted content signal is exactly 0.000 on all twelve items implies the whole measured 'preference' turning out to be display position is a critical way to view elicitation as measurement bias.

    Pre-registered H1 and H2 and reported both as rejected, bootstrapped over items instead of API calls. Counterbalanced exactly in Studies 2b and 3. And declined to claim equivalence in Study 3 because you hadn't pre-registered a margin.

    One thing: Your headline convergence number (Gemini, rho = +0.868) is the one result you never ran through your own position/content decomposition, and it comes from the study whose randomised design left 22.7% of trade-off cells without a matched order. And high cross-method agreement is equally consistent with a shared response heuristic, so you need a discriminant check for convergence.

Cite this project

@misc{rehman2026preflens,
  title = {{PrefLens: Position Artefacts and Convergent Validity in LLM Preference Elicitation}},
  author = {Hafiza Hajrah Rehman and Ameesha},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/preflens-position-artefacts-and-convergent-validity-in-llm-preference-elicitation-hg4c}},
  url = {https://apartresearch.com/sprints/projects/preflens-position-artefacts-and-convergent-validity-in-llm-preference-elicitation-hg4c}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026