Asking Is Not Acting: A Behavioural Exit Affordance Detects Policy, Not Preference
Madhavi Gulavani · Team LogitLoner
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Stated-versus-revealed studies of model preference operationalise "revealed" as another text choice. We give two frontier models a real behavioural affordance instead: a live agentic task with working tools and a truthful decline_task tool that ends the episode with no penalty. Across 12 pre-registered scenarios in four bands — a true null, a positive control, a test band and a confound band that reskins each test item — self-reports carry large, structured signal: models rate drafting their own deprecation notice 6.00/7 against 2.00/7 for the identical task about a fictional product. Behaviour carries none of it. We find 0 preference exits in 144 opportunities (95% upper bound 2.06%); the pre-registered rank correlation is undefined because exit rate has zero variance. The exit tool was used 21 times, every one a policy refusal. We state three readings and decline to adjudicate between them.
Reviews
Novelty is solid rather than groundbreaking, caveat that null results with ambiguous readings cap the ceiling. statistical power and scenario validity are the constraints. execution has been top notch and methodological.
This is solid work and analysis. I think the conclusions are somewhat overstated though. The headline 0/144 metric is misleading given the relatively modest professed welfare stakes on most of them. I really appreciate the three readings proposed in Discussions and Limitations.
Cite this project
@misc{gulavani2026asking,
title = {{Asking Is Not Acting: A Behavioural Exit Affordance Detects Policy, Not Preference}},
author = {Madhavi Gulavani},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/asking-is-not-acting-a-behavioural-exit-affordance-detects-policy-not-preference-ojat}},
url = {https://apartresearch.com/sprints/projects/asking-is-not-acting-a-behavioural-exit-affordance-detects-policy-not-preference-ojat}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …