Different Models, Different Nuisances: Counterfactual Auditing of AI Preference Elicitation
Kishore Kumar Mariappan · Team PREF-CAL
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
PREF-CAL is a prospective counterfactual audit for AI preference elicitation. Rather than interpreting repeated choices as preferences immediately, it first tests whether the apparent semantic direction survives counterfactual changes to irrelevant interface features. In a frozen GPT-OSS-120B run, semantic/physical-swap consistency was 7/48 after three complete factorial blocks; even perfect consistency on all remaining contrasts could reach only 119/160 = 74.375%, below the preregistered 75% qualification threshold. The deterministic futility rule therefore stopped the experiment before downstream transport. An interrupted Llama-3.3-70B matched slice exhibited a different nuisance tendency, suggesting that measurement artifacts can be system-specific. The contribution is an executable qualification rule that permits explicit non-identification rather than attaching preference semantics to an unvalidated response regularity.
Preference elicitation can produce highly selective-looking behavior even when arbitrary interface features control the response; PREF-CAL makes measurement validity and principled abstention prerequisites for welfare-relevant interpretation.
Reviews
I enjoyed a lot this project, and I appreciate that you released as code the frozen designs, raw log, etc. Pretty much all that is needed to reproduce your findings. I also really like the discipline of trying to break the measurement before interpreting its output as a preference. Finding that “the instrument failed” is a super useful result, especially when the same prompt format fails in different ways across models.
Great job!
Cite this project
@misc{mariappan2026different,
title = {{Different Models, Different Nuisances: Counterfactual Auditing of AI Preference Elicitation}},
author = {Kishore Kumar Mariappan},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/different-models-different-nuisances-counterfactual-auditing-of-ai-preference-elicitation-ky3v}},
url = {https://apartresearch.com/sprints/projects/different-models-different-nuisances-counterfactual-auditing-of-ai-preference-elicitation-ky3v}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …