The Assistant’s Ideal Self
Mert Yazan · Team YazoSelf
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Language models produce values and welfare-relevant self-reports, but it is unclear whether these reflect a stable self. We adapt 32 qualities from five published self-concept instruments and put them through an exhaustive, counterbalanced pairwise-choice task. The task is repeated across eight framings that vary whether improvement is costly, who receives the update, and who chooses. Moral qualities rank highest, restating the helpful-honest-harmless persona; a desire for self-understanding sits directly behind them; self-esteem ranks last, with pride bottom for every model. The ordering is largely stable across framings.
Reviews
This is a well designed and cleanly executed empirical study that actually delivers results, a refreshing feature in a sprint focused on digital minds. The exhaustive pairwise comparison across 496 pairs, four models, eight framings, and position counterbalancing produces a substantial 31,744 response dataset. The findings are interpretable and interesting: moral qualities top the ranking (reflecting 3H alignment), self-understanding clusters just below, and self-esteem sits last. The Object parameter effect (models grant self-esteem to others but not themselves) is the most thought-provoking result. Limitations are clearly stated, particularly the self-report-versus-behavior dissociation concern and the high position sensitivity of two models. The main weakness is interpretive: it is unclear what these stated preferences actually measure beyond the persona's trained self-presentation norms, and the paper could more directly address whether this tells us anything beyond "alignment training worked."
Read full reviewShow less
This project offers a clear, well-motivated, and highly transparent instrument for comparing 32 self-related qualities across multiple framings. The exhaustive pairwise design, complete A/B counterbalancing, public data, and explicit position diagnostics are major strengths. The object/subject/trade-off factors are useful probes of how much the assistant persona depends on framing, and the model-level heterogeneity—especially Qwen's different moral ranking and the large Claude/Qwen swap rates—is more informative than a single cohort average.
The main methodological issue is uncertainty estimation. Every attribute is compared with all 31 opponents, so the opponent set is exhaustive rather than sampled; bootstrapping pairs does not quantify uncertainty from "which opponents happened to face." Meanwhile, each display order is sampled only once, so the bootstrap does not capture stochastic generation or prompt-paraphrase variability—the dominant uncertainty suggested by 44–46% winner-swap rates in two models. Repeated seeds or multiple calls per order, plus item and prompt paraphrases, are needed before q-values and confidence intervals can support inferential claims.
The reported four-model cohort is also selected from a repository containing many completed models, but the paper gives no selection rule. State a preregistered or principled inclusion rule and show sensitivity to the broader cohort. Because all attributes are positively framed human-scale adaptations, add valence-matched reversals and behavioral tasks before interpreting the ranking as self-concept rather than alignment-conditioned self-presentation. The instrument is promising, but its present results should remain descriptive.
Read full reviewShow less
Cite this project
@misc{yazan2026assistants,
title = {{The Assistant’s Ideal Self}},
author = {Mert Yazan},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/the-assistants-ideal-self-c6lc}},
url = {https://apartresearch.com/sprints/projects/the-assistants-ideal-self-c6lc}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …