Measuring Preference Coherence, Risk Sensitivity, and Expected Utility Trade-offs in Large Language Models
Aryan masani
This study evaluates how system prompt framings alter the internal consistency, risk sensitivity, and economic decision-making of large language models (LLMs) across 10 financial and operational scenarios. Core Findings Expected Value Maximization: In unconstrained default framings, models act primarily as expected monetary value (EV) maximizers, selecting higher-yielding options or taking gambles in loss domains to maximize overall expected net payoff. Constraint-Driven Risk Aversion: Introducing strict budget or risk-minimization constraints causes models to abandon pure EV maximization, prioritizing baseline survival and downside protection over higher-yield options with severe tail risks. Liquidity Trade-offs: Under resource constraints, models shift toward preserving immediate liquidity (e.g., selecting monthly subscriptions) rather than minimizing long-term cumulative outlays (e.g., lifetime purchases). Utility Reweighting: System prompt directives act as systematic re weightings of a model's internal utility calculation—introducing variance penalties—rather than creating random decision noise.
The design is well matched to the question because it separates risk preference from differences in expected payoff. The gain/loss mirror is also a sensible reflection-effect probe, and the scenarios are varied and thoughtfully constructed. Both reported effects point in the direction predicted by the behavioral literature. The main weakness is not the stimuli but the measurement protocol.
Some things you could push for:
1. Identify the model and inference setup. Report the exact model or snapshot, decoding settings, and run date. The manuscript currently makes claims about “models” without specifying what system produced the data, which makes the result difficult to reproduce or interpret.
2. Sample each cell repeatedly and report choice proportions rather than a single label. With one draw per condition, a stable preference and an approximately 50/50 response look the same in the table. Repeated trials would turn each cell into an estimate with visible uncertainty.
3. Counterbalance option order. The safer option appears first in every reported row, so the gain-domain result could partly reflect a first-position bias. Re-run each item with the options swapped and combine the two orders.
Two claims also need correction. The paper reports strong transitivity, but the current design uses isolated pairwise choices from different scenarios. Because there is no shared set of three or more alternatives, the data do not support a standard transitivity or cycle analysis. Either add choices from common triads and test for cycles, or remove that claim. The Results also mention a Consultant framing that does not appear in the reported table.
The scenario set is a strong foundation. A natural extension would be to test whether different system-prompt personas shift revealed risk preference using the same equal-value pairs, once the basic replication and order controls are in place.
Cite this work
@misc {
title={
(HckPrj) Measuring Preference Coherence, Risk Sensitivity, and Expected Utility Trade-offs in Large Language Models
},
author={
Aryan masani
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


