When Agreement Misleads: Detecting Correlated Errors in Multi-Method LLM Preference Elicitation
Noah Moran
We implement 3 elicitation methods on Qwen3-8B-Instruct on 10 AI welfare outcomes based on: direct choice, avoidance [3], and choosing a gamble. We reduce the results of each method to a directional measure, to avoid faulty comparisons between incomparable scales.
Separately, 66 anchors of varying difficulty were created to validate each method and measure if their errors are correlated. We find that the avoidance elicitation is confidently incorrect 65% of the time, and akin to inverting direct choice. Additionally, we find that direct choice and the gamble agree 89% of the time and their joint failures occur 4.2x more often than independence predicts. All errors are in the hard difficulty, with borderline statistical significance of correlation (p=.08).
The core question — do convergent preference-elicitation methods actually provide independent evidence, or can they share a correlated failure mode the way LLM judge panels do — is a genuinely useful one to ask, and applying Kohli's recent judge-panel finding to elicitation methods rather than judge models is a clean, novel transfer. The two headline findings differ sharply in how well-supported they are, and this should be weighted accordingly.
The avoidance-method finding is the stronger result: method B (avoidance/veto) is confidently wrong 65% of the time and behaves almost as an inversion of direct choice — a clear, actionable finding for anyone using avoidance-framed elicitation as a preference-measurement tool.
The second finding — that direct choice and the gambling method share a correlated failure mode (joint failure 4.2x more common than independence predicts) — is the paper's more novel and more safety-relevant claim, but it rests on a very small evidence base (54 data points after excluding abstentions) and the authors' own reported p-value (.077) does not clear conventional significance. The paper is commendably explicit about this ("statistically inconclusive," "suggests larger datasets should be tested") rather than overstating it, which is the right call — but it does mean this central claim should be read as a promising signal to follow up on, not an established finding.
Two presentation issues : the worked examples for the "Medium" and "Hard" difficulty anchors are printed as identical text, which as written makes the two tiers indistinguishable to a reader (the prose describes a distinction that the examples don't show); and reference [5] misspells the cited author's surname (Shi, not Shin). Neither affects the validity of the underlying experiment, but both should be corrected.
The study is single-model (Qwen3-8B-Instruct only, chosen specifically because it makes enough errors to measure correlation — a reasonable and transparently-stated design choice, though it does mean nothing here speaks to whether the same correlated-failure pattern holds in more capable models where errors are rarer).
Measuring how implicit cues override explicit numerical objectives is an interesting and informative approach. Correlated overrides across elicitations suggest that those elicitations should be treated as highly correlated evidence. But mapping this approach to welfare questions without any ground truth labels available seems fraught, and this submission doesn't report enough welfare results to suggest a significant update. Overall in its current form this approach offers only weak evidence about whether convergent welfare reports reflect genuine preferences.
Cite this work
@misc {
title={
(HckPrj) When Agreement Misleads: Detecting Correlated Errors in Multi-Method LLM Preference Elicitation
},
author={
Noah Moran
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


