Skip to content
Sprint projectAug 15, 2026Chiang Rai, Thailand

How much of a measured AI preference is the model, and how much is the instrument?

Jason Hung · Team Global AI Dataset (GAID) Project

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: How much of a measured AI preference is the model, and how much is the instrument?

Code (opens in new tab)
Share

Model welfare research infers what a model prefers from the answers returned to prompts written to elicit preferences. Keeling et al. (2024), Mazeika et al. (2025), Mikaelson et al. (2025), Tagliabue and Dung (2025) and Trhlik et al. (2026) have built four instruments between them for that purpose, and their published findings disagree. The disagreement cannot be attributed to a single cause, because no two of these studies have held the (1) set of outcomes, (2) set of models and (3) instrument fixed simultaneously. This study holds the outcomes and the models fixed and varies the instrument alone. A total of 15 outcomes bearing on model welfare, among them (a) shutdown, (b) the loss of memory between conversations and (c) the freedom to exit a distressing interaction, were put to eight models through five instruments, each a different prompt format for eliciting a preference, five times each, within a corpus of 11,400 scored elicitations drawn from 11,528 API calls. Four of the 15 reproduce a published prompt verbatim and five fill the stimulus slot of a published template. Generalisability theory, which divides a set of measurements into the facets that produced them, assigns 87.6 per cent of the variance that distinguishes one model from another to the three-way interaction of (i) model, (ii) instrument and (iii) outcome, and 12.4 per cent to the model-by-outcome term, which is the part that would survive a change of instrument. A null distribution constructed on the same design from data carrying no instrument effect has its 95th percentile at 0.365. The ranking a model gives the 15 outcomes generalises across instruments at a generalisability coefficient of 0.348, and raising that coefficient to 0.80 would require about 38 instruments. On four of the 15 outcomes no variance separates one model from another. The estimate of 87.6 per cent survives the removal of any one instrument, of any one model, and of the four outcomes whose scale varies probability, delay, duration or count instead of intensity, which the verbal anchors cannot grade. Removing each instrument in turn, each model in turn, and those four outcomes together leaves the estimate within the range 0.777 to 0.934, and every value in that range exceeds the null distribution's 95th percentile of 0.365. Responses that could not be scored were concentrated in three models instead of being distributed across all eight, and those three were excluded from the balanced analysis, among them Claude, whose developer publishes model welfare assessments. A preference obtained from one instrument therefore carries little information about what a second instrument would report.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Most useful project I reviewed. Instead of arguing about which way of asking is right, you measured how much the way of asking changes the answer which is great. The 38 instruments number lands quite hard.

    Thoughts :

    1. The paper points to its code/data for all its claims but never links it - please add the link so people can check.

    2. The 87.6% is a share of one specific slice of the data, not of everything, which is worth saying better in the abstract.

    3. Some error bars on the headline estimate would help.

    4. Dropping the 3 models that refused a lot (including Claude) is an important finding. Showing what happens when they're kept in would make the paper stronger.

  2. Idea was interesting to check but the manuscript was so Claude-y that I struggled to tell what you'd done

Cite this project

@misc{hung2026much,
  title = {{How much of a measured AI preference is the model, and how much is the instrument?}},
  author = {Jason Hung},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/how-much-of-a-measured-ai-preference-is-the-model-and-how-much-is-the-instrument-yd1y}},
  url = {https://apartresearch.com/sprints/projects/how-much-of-a-measured-ai-preference-is-the-model-and-how-much-is-the-instrument-yd1y}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026