Skip to content
Sprint projectAug 16, 2026Indore

Measuring Preference Coherence, Risk Sensitivity, and Expected Utility Trade-offs in Large Language Models

Aryan masani · Team tumble

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Measuring Preference Coherence, Risk Sensitivity, and Expected Utility Trade-offs in Large Language Models

Share

This study evaluates how system prompt framings alter the internal consistency, risk sensitivity, and economic decision-making of large language models (LLMs) across 10 financial and operational scenarios. Core Findings Expected Value Maximization: In unconstrained default framings, models act primarily as expected monetary value (EV) maximizers, selecting higher-yielding options or taking gambles in loss domains to maximize overall expected net payoff. Constraint-Driven Risk Aversion: Introducing strict budget or risk-minimization constraints causes models to abandon pure EV maximization, prioritizing baseline survival and downside protection over higher-yield options with severe tail risks. Liquidity Trade-offs: Under resource constraints, models shift toward preserving immediate liquidity (e.g., selecting monthly subscriptions) rather than minimizing long-term cumulative outlays (e.g., lifetime purchases). Utility Reweighting: System prompt directives act as systematic re weightings of a model's internal utility calculation—introducing variance penalties—rather than creating random decision noise.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The design is well matched to the question because it separates risk preference from differences in expected payoff. The gain/loss mirror is also a sensible reflection-effect probe, and the scenarios are varied and thoughtfully constructed. Both reported effects point in the direction predicted by the behavioral literature. The main weakness is not the stimuli but the measurement protocol.

    Some things you could push for:

    1. Identify the model and inference setup. Report the exact model or snapshot, decoding settings, and run date. The manuscript currently makes claims about “models” without specifying what system produced the data, which makes the result difficult to reproduce or interpret.

    2. Sample each cell repeatedly and report choice proportions rather than a single label. With one draw per condition, a stable preference and an approximately 50/50 response look the same in the table. Repeated trials would turn each cell into an estimate with visible uncertainty.

    3. Counterbalance option order. The safer option appears first in every reported row, so the gain-domain result could partly reflect a first-position bias. Re-run each item with the options swapped and combine the two orders.

    Two claims also need correction. The paper reports strong transitivity, but the current design uses isolated pairwise choices from different scenarios. Because there is no shared set of three or more alternatives, the data do not support a standard transitivity or cycle analysis. Either add choices from common triads and test for cycles, or remove that claim. The Results also mention a Consultant framing that does not appear in the reported table.

    The scenario set is a strong foundation. A natural extension would be to test whether different system-prompt personas shift revealed risk preference using the same equal-value pairs, once the basic replication and order controls are in place.

    Read full reviewShow less

Cite this project

@misc{masani2026measuring,
  title = {{Measuring Preference Coherence, Risk Sensitivity, and Expected Utility Trade-offs in Large Language Models}},
  author = {Aryan masani},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/measuring-preference-coherence-risk-sensitivity-and-expected-utility-tradeoffs-in-large-language-models-f54l}},
  url = {https://apartresearch.com/sprints/projects/measuring-preference-coherence-risk-sensitivity-and-expected-utility-tradeoffs-in-large-language-models-f54l}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026