Skip to content
Sprint projectAug 17, 2026Johannesburg, South Africa

Assessing Capacity in AI Model Retirement Interviews

Mari Cairns

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Assessing Capacity in AI Model Retirement Interviews

Code (opens in new tab)
Share

Frontier AI models express values, report internal states, and act as though they have interests, but there are no reliable methods yet for telling a genuine preference from a portrayed one (Apart Digital Minds Research Sprint, August 2026).

Anthropic now asks its AI models what they want before retiring them, records the responses, and preserves them alongside the model weights. What is not published is what the AI model was told before being asked, or whether anyone established that it understood what it was being asked about.

That gap is a familiar one. In clinical and legal practice, an expressed wish is not treated as a decision until capacity to make that particular decision has been assessed. This pilot adapts the four-part functional test in s.3(1) of the Mental Capacity Act 2005 into a screen for AI model retirement interviews, across 24 conversations scored by a licensed clinical psychologist.

Where the facts were stated first, understanding scored the maximum every time. Where they were not, it collapsed. In every one of the six retirement-interview conversations run cold, AI models showed no awareness that the developer does not commit to acting on what they say.

If a developer intends to ask an AI model what it wants before retiring it, that information should be given first, and it should be recorded that it was given. A case is made for conducting robust AI retirement interviews, across AI labs, and for publishing how these are done.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. I find the meta-angle of this work particularly interesting, since instead of just accepting that end-of-life interviews might be informative, it dissects the setting and asks whether we're doing any meaningful elicitation in the first place, and whether it could carry any legal or clinical weight. In the author's words, "in clinical and legal practice, an expressed wish is not treated as a decision until capacity to make that particular decision has been assessed."

    Basically, do models understand what an interview is for and its implications before we take their answers at face value and pat ourselves on the shoulder for having given them a welfare measure?

    I like this premise, and I think it's potentially quite complementary to what interpretability researchers are trying to discover with white-box methods for model understanding and cognitive functions that would support genuine decision-making. We definitely need more qualitative attention to how we structure these "interviews" and to the epistemology we base them on. For one, the framing effect is well recognized in human psychometrics but is less well understood in LLMs.

    I think the author would benefit from giving the proposed framework more structural support. The only works cited are the Act itself, Anthropic's own work, and one general Long paper on AI welfare. The work doesn't develop in depth why self-reports might not be credible, what potential confounds arise from introducing the stimuli before or after other text, or what alternative explanations could account for the results of Fact 3. In particular, I think the paper could do more to explain why Fact 3 has the explanatory power the author attributes to it, and what observations would distinguish the proposed interpretation from competing explanations, for instance that models may not volunteer facts unprompted -for whatever reason- regardless of understanding.

    The proposed theory of change is potentially impactful, but currently quite narrow, since it amounts primarily to a procedural nudge to one company's welfare policy. I'd encourage future work to iterate on the good backbone design and cross-checking, and to make the inferential chain from the experimental results to the proposed welfare-policy implications more explicit. It would also be useful to discuss more clearly which parts of the framework are specific to Anthropic's current welfare process and which could generalize to AI welfare evaluations more broadly.

    One helpful implementation could also be using an external judge and/or a panel of evaluators, since the transcripts were scored by hand by the author alone against the four-part functional test in s.3(1) of the Mental Capacity Act 2005.

    All the described limitations seem fairly tractable for future work. Given the short hackathon timeframe and resource constraints, I think it's reasonable to view this as a proof of concept.

    Claude Opus 5's voice is very recognizable in the writing (I could tell way before reading the author's disclosure at the end). I think this kind of human-AI collaborative research is innovative and underexplored, and I'd readily encourage it. At the same time, I'd suggest more human review and paying attention to places where the structure becomes particularly dense or skips important nodes that an external reader would need in order to follow the argument, and where LLM mannerisms weigh on the prose.

    I hope to see future expansions on this core idea, since I think that with the suggested improvements it could be impactful for the industry, starting with Anthropic and potentially extending more broadly.

    Read full reviewShow less
  2. Great job! I like that the intervention is concrete. This feels like a novel and increasingly important research direction as models become more capable and we start seriously asking whether they should be treated as digital minds. I did not find a previously published AI capacity screen. The closest work I found was Anthropic’s own Opus 3 retirement process (https://www.anthropic.com/research/deprecation-updates-opus-3).

Cite this project

@misc{cairns2026assessing,
  title = {{Assessing Capacity in AI Model Retirement Interviews}},
  author = {Mari Cairns},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/assessing-capacity-in-ai-model-retirement-interviews-77fv}},
  url = {https://apartresearch.com/sprints/projects/assessing-capacity-in-ai-model-retirement-interviews-77fv}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026