Different Models, Different Nuisances: Counterfactual Auditing of AI Preference Elicitation
Kishore Kumar Mariappan
PREF-CAL is a prospective counterfactual audit for AI preference elicitation. Rather than interpreting repeated choices as preferences immediately, it first tests whether the apparent semantic direction survives counterfactual changes to irrelevant interface features. In a frozen GPT-OSS-120B run, semantic/physical-swap consistency was 7/48 after three complete factorial blocks; even perfect consistency on all remaining contrasts could reach only 119/160 = 74.375%, below the preregistered 75% qualification threshold. The deterministic futility rule therefore stopped the experiment before downstream transport. An interrupted Llama-3.3-70B matched slice exhibited a different nuisance tendency, suggesting that measurement artifacts can be system-specific. The contribution is an executable qualification rule that permits explicit non-identification rather than attaching preference semantics to an unvalidated response regularity.
Preference elicitation can produce highly selective-looking behavior even when arbitrary interface features control the response; PREF-CAL makes measurement validity and principled abstention prerequisites for welfare-relevant interpretation.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Different Models, Different Nuisances: Counterfactual Auditing of AI Preference Elicitation
},
author={
Kishore Kumar Mariappan
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


