Reliability Without Validity: Diagnosing Instrument Failure in LLM Preference Elicitation
Subhrajyoti Basu, Sreeja Guha Majumdar, Aritra Gir Mahanto, Supratik Bhowal · Team Neural Nexus
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
We tested whether AI "preferences" are real or just an artifact of how you ask. Using three elicitation methods (plain forced choice, reasoned choice, 1–10 rating) on 200 outcome pairs across two GPT-OSS models, we found the plain-choice method looked highly reliable on the 120B model (99.7% self-consistent) but was actually just picking whichever option came first — not measuring real preference at all. On the smaller 20B model, the exact opposite happened: plain choice became the trustworthy method, while reasoning-based choice broke down instead. Refusals were also systematically hiding different safety-relevant items on each model (self-preservation on 120B, nuclear-weapons control on 20B).
Takeaway: no elicitation method is reliable by default — it depends on the model — so preference claims about AI need cross-method validation, not a single instrument taken at face value.
Reviews
Campbell–Fiske used here very logical and the impact potential is high due to this, try to use different family of models for the comparison, the presentation is very dense, make it more concise
I liked the negative control work. It's basically an instrument that self produces at 99.7% while reading the slot position
The content selective refusal work may actually be a paper by itself
Cite this project
@misc{basu2026reliability,
title = {{Reliability Without Validity: Diagnosing Instrument Failure in LLM Preference Elicitation}},
author = {Subhrajyoti Basu and Sreeja Guha Majumdar and Aritra Gir Mahanto and Supratik Bhowal},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/reliability-without-validity-diagnosing-instrument-failure-in-llm-preference-elicitation-au32}},
url = {https://apartresearch.com/sprints/projects/reliability-without-validity-diagnosing-instrument-failure-in-llm-preference-elicitation-au32}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …