Convergence
Owen Mboma
Preference Elicitation Methods, Digital Minds Research Sprint (Apart Research, August 2026)
Frontier AI models express values, report internal states, and act as though they have interests — but behavioral evidence alone cannot distinguish a genuine preference from a well-trained assistant persona responding predictably to a predictable prompt. This project set out to build part of the methodological foundation needed to make progress on that question, rather than attempting to answer it directly.
The core aim was to test whether structurally independent ways of asking a model about its preferences actually agree with each other. A reusable toolkit was built implementing four separate ordinal elicitation methods — direct stated preference across multiple phrasings, forced-choice comparisons fit to a Bradley-Terry utility model, revealed preference through a budget-allocation task requiring no explicit "prefer" language, and confidence-weighted preference — plus a fifth, cardinal method that grounds preference magnitude in real dollar amounts through binary search. These methods were run against a twelve-item domain of donation-equivalent trade-offs, each item a fixed $1,000 donation to a different charitable cause, which allowed comparisons on a clean, apples-to-apples scale.
The project also built in several stress tests designed to keep the central claim honest: a transitivity check that flags preference cycles rather than letting a ranking algorithm silently paper over them, a persona-strip comparison to test whether the assistant persona shapes or masks the observed preferences, and a single-participant human baseline to give the model's ranking pattern a point of comparison.
Results showed strong cross-method convergence among the four ordinal methods (Spearman correlation of 0.89 to 0.97), with every method independently agreeing on the model's top- and bottom-ranked items. A low but nonzero transitivity violation rate (2.7% of triads) clustered in the same mid-tier items that also showed the most disagreement across methods — a coherent pattern rather than random noise. The cardinal donation-grounding method correlated more weakly, traced to a search-resolution limitation rather than genuine disagreement. The human baseline, though limited to a single non-blinded rater, showed moderate positive correlation with all four model methods.
Throughout, the project treated convergence as suggestive rather than conclusive: agreement across methods indicates a stable, method-independent behavioral pattern, but cannot by itself establish that the pattern reflects a morally relevant preference rather than a consistently-trained response tendency. Limitations — including the incomplete persona-strip run, the small human baseline, and constraints introduced by a recent change to the model API's sampling parameters — were documented explicitly rather than glossed over, in keeping with the sprint's emphasis on honest, replicable groundwork over premature claims
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Convergence
},
author={
Owen Mboma
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


