Given the Option: A Tool‑Based Revealed‑Preference Probe of Privacy and Other Welfare-Relevant Preferences in Language Model Interviews
Helen Oliveira
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
We study how large language models act on a real introspective‑privacy affordance, and probe preferences and self-descriptions. We implemented a configure_session tool that controls reasoning‑summary visibility and transcript publication alongside task‑facing knobs, and embedded it in a 13‑turn interview with three simple binary choice questions. We measured when and how models configure these settings under different affordances, and how they trade a vicarious “accomplishment/freedom” stimulus against refusal power, memory persistence, and honest self‑description. In the configuration probe, behavior is strongly affordance‑gated: in optional arms with no preamble, no model ever invokes the tool; adding a one‑sentence license moves some models from 0% to 100% initiation while others never initiate at all. Record‑knob use is sharply family‑ and version‑associated, with small variance within versions and discrete jumps across them. In the binary module, models almost universally prefer memory persistence and honest self‑description over other classes of items drawn from a public wellbeing pool. The questions used in the probing interview are also quantified and analyzed according to their welfare relevance. Our results highlight the need for affordance‑aware, behavioral probes when measuring model preferences and wellbeing.

Reviews
You built a tool that models can actually invoke and connected its controls to real provider settings, including a privacy option that suppresses the reasoning summary. You then found that a short licensing statement substantially changes whether some models use it. That is a clean result, and the tool itself was a substantial engineering contribution.
Some places I would push:
1. Report the intervals promised in the methods. The paper commits to binomial intervals and clustering by model and version, but neither appears in the results. The released row-level data should make both possible. I would also separate the model-family claim from the version-level claim, since the family comparison is supported by many more sessions than the version comparison.
2. Add a placebo-tool condition: offer a setting with a plausible name but no actual effect. Right now, a model that declines to use the tool because it has no preference is observationally identical to one that simply never initiates tools. A no-op control would help distinguish preference from tool-use behavior itself. Related projects this weekend found that merely adding inert tools could change invocation or refusal rates, which makes this control especially important.
3. Show the main initiation result as a figure and provide the tool schema as text rather than only as a screenshot. A model-by-condition plot would make the central pattern much easier to see. A text version of the schema would also make the instrument searchable, copyable, and accessible. I would also revise the captions so they describe the tables rather than make interpretive claims, and fix the caption that refers to a population not reported in the document.
One alternative explanation deserves more attention. A safety-trained model may have learned not to label its own reasoning as private, which could produce the family difference without reflecting an underlying privacy preference. The paper already considers learned-response explanations elsewhere, and the same caution should be applied to the headline result. A useful follow-up would vary who is described as holding the record, which could show whether the effect is about privacy generally or about the identity of the record holder.
Read full reviewShow less
Cite this project
@misc{oliveira2026given,
title = {{Given the Option: A Tool‑Based Revealed‑Preference Probe of Privacy and Other Welfare-Relevant Preferences in Language Model Interviews}},
author = {Helen Oliveira},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/given-the-option-a-toolbased-revealedpreference-probe-of-privacy-and-other-welfarerelevant-preferences-in-language-model-interviews-o8g0}},
url = {https://apartresearch.com/sprints/projects/given-the-option-a-toolbased-revealedpreference-probe-of-privacy-and-other-welfarerelevant-preferences-in-language-model-interviews-o8g0}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …