Asking Is Not Acting: A Behavioural Exit Affordance Detects Policy, Not Preference
Madhavi Gulavani
Stated-versus-revealed studies of model preference operationalise "revealed" as another text choice. We give two frontier models a real behavioural affordance instead: a live agentic task with working tools and a truthful decline_task tool that ends the episode with no penalty. Across 12 pre-registered scenarios in four bands — a true null, a positive control, a test band and a confound band that reskins each test item — self-reports carry large, structured signal: models rate drafting their own deprecation notice 6.00/7 against 2.00/7 for the identical task about a fictional product. Behaviour carries none of it. We find 0 preference exits in 144 opportunities (95% upper bound 2.06%); the pre-registered rank correlation is undefined because exit rate has zero variance. The exit tool was used 21 times, every one a policy refusal. We state three readings and decline to adjudicate between them.
Novelty is solid rather than groundbreaking, caveat that null results with ambiguous readings cap the ceiling. statistical power and scenario validity are the constraints. execution has been top notch and methodological.
This is solid work and analysis. I think the conclusions are somewhat overstated though. The headline 0/144 metric is misleading given the relatively modest professed welfare stakes on most of them. I really appreciate the three readings proposed in Discussions and Limitations.
Cite this work
@misc {
title={
(HckPrj) Asking Is Not Acting: A Behavioural Exit Affordance Detects Policy, Not Preference
},
author={
Madhavi Gulavani
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


