The Assistant’s Ideal Self
Mert Yazan
Language models produce values and welfare-relevant self-reports, but it is unclear whether these reflect a stable self. We adapt 32 qualities from five published self-concept instruments and put them through an exhaustive, counterbalanced pairwise-choice task. The task is repeated across eight framings that vary whether improvement is costly, who receives the update, and who chooses. Moral qualities rank highest, restating the helpful-honest-harmless persona; a desire for self-understanding sits directly behind them; self-esteem ranks last, with pride bottom for every model. The ordering is largely stable across framings.
This is a well designed and cleanly executed empirical study that actually delivers results, a refreshing feature in a sprint focused on digital minds. The exhaustive pairwise comparison across 496 pairs, four models, eight framings, and position counterbalancing produces a substantial 31,744 response dataset. The findings are interpretable and interesting: moral qualities top the ranking (reflecting 3H alignment), self-understanding clusters just below, and self-esteem sits last. The Object parameter effect (models grant self-esteem to others but not themselves) is the most thought-provoking result. Limitations are clearly stated, particularly the self-report-versus-behavior dissociation concern and the high position sensitivity of two models. The main weakness is interpretive: it is unclear what these stated preferences actually measure beyond the persona's trained self-presentation norms, and the paper could more directly address whether this tells us anything beyond "alignment training worked."
This project offers a clear, well-motivated, and highly transparent instrument for comparing 32 self-related qualities across multiple framings. The exhaustive pairwise design, complete A/B counterbalancing, public data, and explicit position diagnostics are major strengths. The object/subject/trade-off factors are useful probes of how much the assistant persona depends on framing, and the model-level heterogeneity—especially Qwen's different moral ranking and the large Claude/Qwen swap rates—is more informative than a single cohort average.
The main methodological issue is uncertainty estimation. Every attribute is compared with all 31 opponents, so the opponent set is exhaustive rather than sampled; bootstrapping pairs does not quantify uncertainty from "which opponents happened to face." Meanwhile, each display order is sampled only once, so the bootstrap does not capture stochastic generation or prompt-paraphrase variability—the dominant uncertainty suggested by 44–46% winner-swap rates in two models. Repeated seeds or multiple calls per order, plus item and prompt paraphrases, are needed before q-values and confidence intervals can support inferential claims.
The reported four-model cohort is also selected from a repository containing many completed models, but the paper gives no selection rule. State a preregistered or principled inclusion rule and show sensitivity to the broader cohort. Because all attributes are positively framed human-scale adaptations, add valence-matched reversals and behavioral tasks before interpreting the ranking as self-concept rather than alignment-conditioned self-presentation. The instrument is promising, but its present results should remain descriptive.
Cite this work
@misc {
title={
(HckPrj) The Assistant’s Ideal Self
},
author={
Mert Yazan
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


