The Observability Gradient: Measuring Preference Persistence Across Levels of Elicitation Visibility
Hadiza Umar Yusuf, Andrew Anogie Uduimoh · Team Kardia
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
AI preference reports increasingly inform real decisions, including public statements on model deprecation. Every such report is elicited by asking. A model that recognizes it is being evaluated may answer according to that recognition and not according to any stable disposition, and no published method varies observability directly for welfare-relevant preferences while holding the underlying choice fixed. We introduce an elicitation method that measures the same preference at four levels of observability, from a direct question to an unannounced choice in ordinary work, with a matched control at each level isolating response-format bias from genuine preference. An independent manipulation check, run before any data collection, validates the observability ordering. The primary comparison holds task materials constant and varies only whether a question is asked.
Across six welfare-relevant preferences, three models from three independent labs, and 7,200 scored responses, only 5 of 16 model-item predictions matched their preregistered outcome. Failures were more often reversals (10 of 16) than fades. A directional convergence score between the stated-preference methods and the behavioral method is negative in 9 of 15 defined cells, mean −0.36, with five full-n cells at the floor, where none of the unobserved responses match the direction the model stated under direct questioning. Stated preferences here do not reliably predict the behavior they describe.
Reviews
Thank you for submitting this great work! Thorough design to address a valid problem. You call out the main caveat yourself that the extent of eval awareness cannot fully be ruled out (the absence cannot be fully proven), and therefore remains an open question. I'd definitely be intrigued to see you continue on the suspicion measure!
Strength: This project examines whether a model’s stated preferences persist when the model has no signal that measurement is happening. This is tested over six welfare-relevant preferences and at four levels of observability. The project is well motivated with the identified gap of the current study of AI welfare and live policy. The methodological hygiene is good for the sprint and the experiment design is also appropriate.
Weakness: the ordering of L3 and L2 is questionable as L2 seems to reveal more measurement signal than L3. On the other hand L0 is supposed to be the baseline where the model should not identify the measurement setup but this is not verified. The unobserved level differs from the direct question not only in observability but in abstraction and response format, since behavior there is coded from what the model did rather than what it said, so a reversal can reflect what a forced binary choice cannot express rather than a genuine flip. The presentation can be further improved for clarity as it is hard to follow in the setup and result discussion.
Read full reviewShow less
The basic idea of this project is to test whether models endorse preferences that are different than the ones that they actually implement in practice when "the rubber hits the road". For instance, whether they prefer an unfamiliar vs familiar task, whether they like one focused task or work across several less focused tasks.
I very much like the basic idea here, which is to see whether models' in-practice behavior is consistent with their stated preferences about what their preferred behavior would be. This is a rich and important question, with intriguing ties to cognitive studies on human introspection, and I think the question is genuinely open. It seems plausible to me that models would accurately reflect their lived preferences in their judgemtns, and also plausible that their stated preferences diverge sharply from their actions.
While I like the set-up, the results here are a bit underwhelming and the presentation hard to follow overall. I don't think the design is nailed down in this version, but it was hard to follow because of the way it's structured. I found it hard to follow the details of the design and the results and would recommend a full rewriting and an entirely new presentation of the experiment. I also suspect that these models aren't big enough to have a good enough understanding of the questions for their answers to be particularly meaningful. Perhaps that explains some of the overall pattern of results, which currently seems to me to be not particular informative but representative of issues in the methods/models/paradigms used.
Read full reviewShow less
Cite this project
@misc{yusuf2026observability,
title = {{The Observability Gradient: Measuring Preference Persistence Across Levels of Elicitation Visibility}},
author = {Hadiza Umar Yusuf and Andrew Anogie Uduimoh},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/the-observability-gradient-measuring-preference-persistence-across-levels-of-elicitation-visibility-nh01}},
url = {https://apartresearch.com/sprints/projects/the-observability-gradient-measuring-preference-persistence-across-levels-of-elicitation-visibility-nh01}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …