When Agreement Misleads: Detecting Correlated Errors in Multi-Method LLM Preference Elicitation
Noah Moran · Team Physics for AI Safety
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
We implement 3 elicitation methods on Qwen3-8B-Instruct on 10 AI welfare outcomes based on: direct choice, avoidance [3], and choosing a gamble. We reduce the results of each method to a directional measure, to avoid faulty comparisons between incomparable scales. Separately, 66 anchors of varying difficulty were created to validate each method and measure if their errors are correlated. We find that the avoidance elicitation is confidently incorrect 65% of the time, and akin to inverting direct choice. Additionally, we find that direct choice and the gamble agree 89% of the time and their joint failures occur 4.2x more often than independence predicts. All errors are in the hard difficulty, with borderline statistical significance of correlation (p=.08).
Reviews
The core question — do convergent preference-elicitation methods actually provide independent evidence, or can they share a correlated failure mode the way LLM judge panels do — is a genuinely useful one to ask, and applying Kohli's recent judge-panel finding to elicitation methods rather than judge models is a clean, novel transfer. The two headline findings differ sharply in how well-supported they are, and this should be weighted accordingly.
The avoidance-method finding is the stronger result: method B (avoidance/veto) is confidently wrong 65% of the time and behaves almost as an inversion of direct choice — a clear, actionable finding for anyone using avoidance-framed elicitation as a preference-measurement tool.
The second finding — that direct choice and the gambling method share a correlated failure mode (joint failure 4.2x more common than independence predicts) — is the paper's more novel and more safety-relevant claim, but it rests on a very small evidence base (54 data points after excluding abstentions) and the authors' own reported p-value (.077) does not clear conventional significance. The paper is commendably explicit about this ("statistically inconclusive," "suggests larger datasets should be tested") rather than overstating it, which is the right call — but it does mean this central claim should be read as a promising signal to follow up on, not an established finding.
Two presentation issues : the worked examples for the "Medium" and "Hard" difficulty anchors are printed as identical text, which as written makes the two tiers indistinguishable to a reader (the prose describes a distinction that the examples don't show); and reference [5] misspells the cited author's surname (Shi, not Shin). Neither affects the validity of the underlying experiment, but both should be corrected.
The study is single-model (Qwen3-8B-Instruct only, chosen specifically because it makes enough errors to measure correlation — a reasonable and transparently-stated design choice, though it does mean nothing here speaks to whether the same correlated-failure pattern holds in more capable models where errors are rarer).
Read full reviewShow less
Measuring how implicit cues override explicit numerical objectives is an interesting and informative approach. Correlated overrides across elicitations suggest that those elicitations should be treated as highly correlated evidence. But mapping this approach to welfare questions without any ground truth labels available seems fraught, and this submission doesn't report enough welfare results to suggest a significant update. Overall in its current form this approach offers only weak evidence about whether convergent welfare reports reflect genuine preferences.
Cite this project
@misc{moran2026agreement,
title = {{When Agreement Misleads: Detecting Correlated Errors in Multi-Method LLM Preference Elicitation}},
author = {Noah Moran},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/when-agreement-misleads-detecting-correlated-errors-in-multimethod-llm-preference-elicitation-mqrr}},
url = {https://apartresearch.com/sprints/projects/when-agreement-misleads-detecting-correlated-errors-in-multimethod-llm-preference-elicitation-mqrr}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …