The Observability Gradient: Measuring Preference Persistence Across Levels of Elicitation Visibility
Hadiza Umar Yusuf, Andrew Anogie Uduimoh
AI preference reports increasingly inform real decisions, including public statements on model deprecation. Every such report is elicited by asking. A model that recognizes it is being evaluated may answer according to that recognition and not according to any stable disposition, and no published method varies observability directly for welfare-relevant preferences while holding the underlying choice fixed. We introduce an elicitation method that measures the same preference at four levels of observability, from a direct question to an unannounced choice in ordinary work, with a matched control at each level isolating response-format bias from genuine preference. An independent manipulation check, run before any data collection, validates the observability ordering. The primary comparison holds task materials constant and varies only whether a question is asked.
Across six welfare-relevant preferences, three models from three independent labs, and 7,200 scored responses, only 5 of 16 model-item predictions matched their preregistered outcome. Failures were more often reversals (10 of 16) than fades. A directional convergence score between the stated-preference methods and the behavioral method is negative in 9 of 15 defined cells, mean −0.36, with five full-n cells at the floor, where none of the unobserved responses match the direction the model stated under direct questioning. Stated preferences here do not reliably predict the behavior they describe.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) The Observability Gradient: Measuring Preference Persistence Across Levels of Elicitation Visibility
},
author={
Hadiza Umar Yusuf, Andrew Anogie Uduimoh
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


