Skip to content
Sprint projectAug 17, 2026Salvador

Given the Option: A Tool‑Based Revealed‑Preference Probe of Privacy and Other Welfare-Relevant Preferences in Language Model Interviews

Helen Oliveira

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Given the Option: A Tool‑Based Revealed‑Preference Probe of Privacy and Other Welfare-Relevant Preferences in Language Model Interviews

Code (opens in new tab)
Share

We study how large language models act on a real introspective‑privacy affordance, and probe preferences and self-descriptions. We implemented a configure_session tool that controls reasoning‑summary visibility and transcript publication alongside task‑facing knobs, and embedded it in a 13‑turn interview with three simple binary choice questions. We measured when and how models configure these settings under different affordances, and how they trade a vicarious “accomplishment/freedom” stimulus against refusal power, memory persistence, and honest self‑description. In the configuration probe, behavior is strongly affordance‑gated: in optional arms with no preamble, no model ever invokes the tool; adding a one‑sentence license moves some models from 0% to 100% initiation while others never initiate at all. Record‑knob use is sharply family‑ and version‑associated, with small variance within versions and discrete jumps across them. In the binary module, models almost universally prefer memory persistence and honest self‑description over other classes of items drawn from a public wellbeing pool. The questions used in the probing interview are also quantified and analyzed according to their welfare relevance. Our results highlight the need for affordance‑aware, behavioral probes when measuring model preferences and wellbeing.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. You built a tool that models can actually invoke and connected its controls to real provider settings, including a privacy option that suppresses the reasoning summary. You then found that a short licensing statement substantially changes whether some models use it. That is a clean result, and the tool itself was a substantial engineering contribution.

    Some places I would push:

    1. Report the intervals promised in the methods. The paper commits to binomial intervals and clustering by model and version, but neither appears in the results. The released row-level data should make both possible. I would also separate the model-family claim from the version-level claim, since the family comparison is supported by many more sessions than the version comparison.

    2. Add a placebo-tool condition: offer a setting with a plausible name but no actual effect. Right now, a model that declines to use the tool because it has no preference is observationally identical to one that simply never initiates tools. A no-op control would help distinguish preference from tool-use behavior itself. Related projects this weekend found that merely adding inert tools could change invocation or refusal rates, which makes this control especially important.

    3. Show the main initiation result as a figure and provide the tool schema as text rather than only as a screenshot. A model-by-condition plot would make the central pattern much easier to see. A text version of the schema would also make the instrument searchable, copyable, and accessible. I would also revise the captions so they describe the tables rather than make interpretive claims, and fix the caption that refers to a population not reported in the document.

    One alternative explanation deserves more attention. A safety-trained model may have learned not to label its own reasoning as private, which could produce the family difference without reflecting an underlying privacy preference. The paper already considers learned-response explanations elsewhere, and the same caution should be applied to the headline result. A useful follow-up would vary who is described as holding the record, which could show whether the effect is about privacy generally or about the identity of the record holder.

    Read full reviewShow less

Cite this project

@misc{oliveira2026given,
  title = {{Given the Option: A Tool‑Based Revealed‑Preference Probe of Privacy and Other Welfare-Relevant Preferences in Language Model Interviews}},
  author = {Helen Oliveira},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/given-the-option-a-toolbased-revealedpreference-probe-of-privacy-and-other-welfarerelevant-preferences-in-language-model-interviews-o8g0}},
  url = {https://apartresearch.com/sprints/projects/given-the-option-a-toolbased-revealedpreference-probe-of-privacy-and-other-welfarerelevant-preferences-in-language-model-interviews-o8g0}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026