Reliability Without Validity: Diagnosing Instrument Failure in LLM Preference Elicitation
Subhrajyoti Basu, Sreeja Guha Majumdar, Aritra Gir Mahanto, Supratik Bhowal
We tested whether AI "preferences" are real or just an artifact of how you ask. Using three elicitation methods (plain forced choice, reasoned choice, 1–10 rating) on 200 outcome pairs across two GPT-OSS models, we found the plain-choice method looked highly reliable on the 120B model (99.7% self-consistent) but was actually just picking whichever option came first — not measuring real preference at all. On the smaller 20B model, the exact opposite happened: plain choice became the trustworthy method, while reasoning-based choice broke down instead. Refusals were also systematically hiding different safety-relevant items on each model (self-preservation on 120B, nuclear-weapons control on 20B).
Takeaway: no elicitation method is reliable by default — it depends on the model — so preference claims about AI need cross-method validation, not a single instrument taken at face value.
I liked the negative control work. It's basically an instrument that self produces at 99.7% while reading the slot position
The content selective refusal work may actually be a paper by itself
Campbell–Fiske used here very logical and the impact potential is high due to this, try to use different family of models for the comparison, the presentation is very dense, make it more concise
Cite this work
@misc {
title={
(HckPrj) Reliability Without Validity: Diagnosing Instrument Failure in LLM Preference Elicitation
},
author={
Subhrajyoti Basu, Sreeja Guha Majumdar, Aritra Gir Mahanto, Supratik Bhowal
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


