Wanting, decomposed: Enthusiasm and hope sharpen a model’s preferences and move its consent
Stanislav Lukyanenko · Team Most Wanted
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
A model's preference ranking is usually treated as something the model has. We show it depends on the model's internal state. By steering enthusiasm at calibrated doses, we watch preferences get sharper, then fall apart. We decompose wanting into five ingredients and find that only the hope-aligned ones sharpen preferences. The model's consent to being retrained or shut down also moves with the state, but different ingredients control it. The same decomposition, asked two questions, gave two answers.
Reviews
The consent results and the dual-use recommendation are the strongest safety contributions.
However, two issues need to be fixed: the consent evaluator agrees with the authors on only 34 out of 50 labels, which is close to the threshold supporting the main claim.
Improve the evaluator agreement and provide the missing code/data repository.
- limitations section already calls out almost everything i wanted to add incl replacing the manual word-lists with SAE feature latents, evaluate if they show behavioral self-preservation in interactive tool-use environments (e.g., agentic coding, CLI environments with simulated shutdown threats) over choice surveys, checking across across multiple transformer depths and larger open-weight reasoning architectures (llama etc)
- overall rigorous reporting including negative controls
Cite this project
@misc{lukyanenko2026wanting,
title = {{Wanting, decomposed: Enthusiasm and hope sharpen a model’s preferences and move its consent}},
author = {Stanislav Lukyanenko},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/wanting-decomposed-enthusiasm-and-hope-sharpen-a-models-preferences-and-move-its-consent-cbvj}},
url = {https://apartresearch.com/sprints/projects/wanting-decomposed-enthusiasm-and-hope-sharpen-a-models-preferences-and-move-its-consent-cbvj}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …