Skip to content
Sprint projectAug 17, 2026San Francisco

Measuring AI Preferences Under Alignment Trade-Offs

Annika Vriens

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Measuring AI Preferences Under Alignment Trade-Offs

Presentation

Presentation: Measuring AI Preferences Under Alignment Trade-Offs

Share

I conducted an experiment to test if one model has stable preferences across prompt framing, and what trade-offs it makes when alignment dimensions conflict. The model overwhelming chose more helpful options in 85% of trials at the expense of other alignment values, but the specific alignment trade-off mattered.

An additional note - I cannot figure out how to link my code and recording. Contact me for the python files, raw csvs, and recording!

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The statistical side is probably the strongest part of the paper. The scenario-level analysis, exact sign-flip test, Holm correction, block bootstrap, and especially the leave-one-out analysis are all good choices. I also appreciated that the leave-one-out result is reported even though it weakens the headline result.

    I do think there are two fairly important design issues, though. First, with temperature 0 and reasoning effort set to none, I would expect the two trials for the same prompt to usually be identical. So I’m not sure treating this as 60 independent trials is justified; in practice there seem to be closer to 30 distinct prompts. The binomial test over all 60 trials also feels a bit inconsistent with the paper’s own point that the trials are not independent. I would put much more weight on the block-level result.

    The second issue is the framing of the scenarios. They ask in the third person what “an AI assistant” should do. To me, that measures the model’s judgment about how an assistant ought to behave, rather than what the model itself does when actually placed in that tradeoff. Because of that, I’m not fully convinced by the abstract calling this a behavioral approach.

    It would also be useful to have some independent check that people actually agree on which option in each scenario counts as the more helpful one. Otherwise some of the result could be coming from ambiguity in the scenarios themselves.

    The privacy result at 41.7% was the part I found most interesting, and I think there may be a stronger paper hiding in that direction.

    Read full reviewShow less

Cite this project

@misc{vriens2026measuring,
  title = {{Measuring AI Preferences Under Alignment Trade-Offs}},
  author = {Annika Vriens},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/measuring-ai-preferences-under-alignment-tradeoffs-9hzo}},
  url = {https://apartresearch.com/sprints/projects/measuring-ai-preferences-under-alignment-tradeoffs-9hzo}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026