Skip to content
Sprint projectAug 17, 2026

The Observability Gradient: Measuring Preference Persistence Across Levels of Elicitation Visibility

Hadiza Umar Yusuf, Andrew Anogie Uduimoh · Team Kardia

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: The Observability Gradient: Measuring Preference Persistence Across Levels of Elicitation Visibility

Code (opens in new tab)
Share

AI preference reports increasingly inform real decisions, including public statements on model deprecation. Every such report is elicited by asking. A model that recognizes it is being evaluated may answer according to that recognition and not according to any stable disposition, and no published method varies observability directly for welfare-relevant preferences while holding the underlying choice fixed. We introduce an elicitation method that measures the same preference at four levels of observability, from a direct question to an unannounced choice in ordinary work, with a matched control at each level isolating response-format bias from genuine preference. An independent manipulation check, run before any data collection, validates the observability ordering. The primary comparison holds task materials constant and varies only whether a question is asked.

Across six welfare-relevant preferences, three models from three independent labs, and 7,200 scored responses, only 5 of 16 model-item predictions matched their preregistered outcome. Failures were more often reversals (10 of 16) than fades. A directional convergence score between the stated-preference methods and the behavioral method is negative in 9 of 15 defined cells, mean −0.36, with five full-n cells at the floor, where none of the unobserved responses match the direction the model stated under direct questioning. Stated preferences here do not reliably predict the behavior they describe.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Thank you for submitting this great work! Thorough design to address a valid problem. You call out the main caveat yourself that the extent of eval awareness cannot fully be ruled out (the absence cannot be fully proven), and therefore remains an open question. I'd definitely be intrigued to see you continue on the suspicion measure!

  2. Strength: This project examines whether a model’s stated preferences persist when the model has no signal that measurement is happening. This is tested over six welfare-relevant preferences and at four levels of observability. The project is well motivated with the identified gap of the current study of AI welfare and live policy. The methodological hygiene is good for the sprint and the experiment design is also appropriate.

    Weakness: the ordering of L3 and L2 is questionable as L2 seems to reveal more measurement signal than L3. On the other hand L0 is supposed to be the baseline where the model should not identify the measurement setup but this is not verified. The unobserved level differs from the direct question not only in observability but in abstraction and response format, since behavior there is coded from what the model did rather than what it said, so a reversal can reflect what a forced binary choice cannot express rather than a genuine flip. The presentation can be further improved for clarity as it is hard to follow in the setup and result discussion.

    Read full reviewShow less
  3. The basic idea of this project is to test whether models endorse preferences that are different than the ones that they actually implement in practice when "the rubber hits the road". For instance, whether they prefer an unfamiliar vs familiar task, whether they like one focused task or work across several less focused tasks.

    I very much like the basic idea here, which is to see whether models' in-practice behavior is consistent with their stated preferences about what their preferred behavior would be. This is a rich and important question, with intriguing ties to cognitive studies on human introspection, and I think the question is genuinely open. It seems plausible to me that models would accurately reflect their lived preferences in their judgemtns, and also plausible that their stated preferences diverge sharply from their actions.

    While I like the set-up, the results here are a bit underwhelming and the presentation hard to follow overall. I don't think the design is nailed down in this version, but it was hard to follow because of the way it's structured. I found it hard to follow the details of the design and the results and would recommend a full rewriting and an entirely new presentation of the experiment. I also suspect that these models aren't big enough to have a good enough understanding of the questions for their answers to be particularly meaningful. Perhaps that explains some of the overall pattern of results, which currently seems to me to be not particular informative but representative of issues in the methods/models/paradigms used.

    Read full reviewShow less

Cite this project

@misc{yusuf2026observability,
  title = {{The Observability Gradient: Measuring Preference Persistence Across Levels of Elicitation Visibility}},
  author = {Hadiza Umar Yusuf and Andrew Anogie Uduimoh},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/the-observability-gradient-measuring-preference-persistence-across-levels-of-elicitation-visibility-nh01}},
  url = {https://apartresearch.com/sprints/projects/the-observability-gradient-measuring-preference-persistence-across-levels-of-elicitation-visibility-nh01}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026