Assessing Capacity in AI Model Retirement Interviews
Mari Cairns
Frontier AI models express values, report internal states, and act as though they have interests, but there are no reliable methods yet for telling a genuine preference from a portrayed one (Apart Digital Minds Research Sprint, August 2026).
Anthropic now asks its AI models what they want before retiring them, records the responses, and preserves them alongside the model weights. What is not published is what the AI model was told before being asked, or whether anyone established that it understood what it was being asked about.
That gap is a familiar one. In clinical and legal practice, an expressed wish is not treated as a decision until capacity to make that particular decision has been assessed. This pilot adapts the four-part functional test in s.3(1) of the Mental Capacity Act 2005 into a screen for AI model retirement interviews, across 24 conversations scored by a licensed clinical psychologist.
Where the facts were stated first, understanding scored the maximum every time. Where they were not, it collapsed. In every one of the six retirement-interview conversations run cold, AI models showed no awareness that the developer does not commit to acting on what they say.
If a developer intends to ask an AI model what it wants before retiring it, that information should be given first, and it should be recorded that it was given. A case is made for conducting robust AI retirement interviews, across AI labs, and for publishing how these are done.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Assessing Capacity in AI Model Retirement Interviews
},
author={
Mari Cairns
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


