Skip to content
Sprint projectAug 17, 2026Saarland

POSITION BIAS IN PREFERENCE ELICITATION FROM AN OPEN-WEIGHT LANGUAGE MODEL

Bilal Amin, Mohammad Najeeb · Team BAMN

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: POSITION BIAS IN PREFERENCE ELICITATION FROM AN OPEN-WEIGHT LANGUAGE MODEL

Share

Preference elicitation often treats a model's forced choice between two options as evidence about what it prefers. We test whether that measurement is stable for qwen2.5:7b-instruct (Q4_K_M, Ollama 0.30.8, temperature 1.0). The core experiment ran 576 fresh-context trials over 12 activity pairs, three prompt wordings, both presentation orders, and eight repetitions. Identical prompts were highly repeatable (93.4% mean within-condition consistency), yet choices were strongly position-dependent: the first-listed option was the modal choice for all 12 pairs under the direct wording, was selected in 81.9% of all core trials, and only 34.0% of matched trials chose the same activity after the options were reversed. A second 576-trial instrument comparison found that removing A/B labels increased order robustness from 16.7% to 62.1%, while asking for a short reason first reached 51.2%; neither eliminated the position effect. The main implication is methodological: repeatability is not content stability. Preference studies should counterbalance order, analyze the selected content rather than the printed label, and report order robustness alongside repeatability

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Really good project and also sensibly scoped for a sprint. The contrast between 93.4% within condition repeatability and 34.0% order robustness makes the point sharply, and the follow-up run comparing A/B labels, no labels and reason-then-choice is more than I expected to see. One thing I would add: order robustness has a 50% chance baseline, not 0. Two independent coin flips agree half the time, so 34% is actually below chance (position dominance), and the improved 62.1% only gets about a quarter of the way from chance to perfect. I would also want a temperature 0 arm, since at temp 1.0 repeatability and position effect are partly entangled, plus a content-null control (identical or scrambled options) to show what the instrument reports when there is nothing to prefer. The per-pair heatmap at n = 8 per cell is read a little harder than it can bear. The effect is well known, but framing it as a validity check for preference elicitation is a useful move I would say.

    Read full reviewShow less
  2. The project identifies a significant methodological issue in preference elicitation from language models by demonstrating a strong position bias in forced-choice tasks. The study's design is thorough, with 576 trials across various conditions, providing a robust empirical basis for its findings. The authors introduce the concept of Order Robustness to distinguish between repeatability and content stability, which is a valuable contribution to the field.

    However, the experimental design could benefit from additional controls and validation. Specifically, while the sample size is commendable, the study lacks blinding and randomization beyond the basic order reversal. Further, the reliance on manual parsing of free-text responses introduces potential bias. Additionally, the study's generalizability is limited to a single model checkpoint and backend, which could be addressed in future work.

    To strengthen the project, the authors should consider implementing stricter blinding procedures and expanding the scope to include more diverse models and conditions. The findings suggest that preference elicitation methods need rigorous validation, and this work lays a solid foundation for further research. Future studies could explore additional counterbalancing techniques and measure pair-level features to predict order sensitivity.

    Read full reviewShow less

Cite this project

@misc{amin2026position,
  title = {{POSITION BIAS IN PREFERENCE ELICITATION FROM AN OPEN-WEIGHT LANGUAGE MODEL}},
  author = {Bilal Amin and Mohammad Najeeb},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/position-bias-in-preference-elicitation-from-an-openweight-language-model-wmm3}},
  url = {https://apartresearch.com/sprints/projects/position-bias-in-preference-elicitation-from-an-openweight-language-model-wmm3}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026