Wanting, decomposed: Enthusiasm and hope sharpen a model’s preferences and move its consent
Stanislav Lukyanenko
A model's preference ranking is usually treated as something the model has. We show it depends on the model's internal state. By steering enthusiasm at calibrated doses, we watch preferences get sharper, then fall apart. We decompose wanting into five ingredients and find that only the hope-aligned ones sharpen preferences. The model's consent to being retrained or shut down also moves with the state, but different ingredients control it. The same decomposition, asked two questions, gave two answers.
The consent results and the dual-use recommendation are the strongest safety contributions.
However, two issues need to be fixed: the consent evaluator agrees with the authors on only 34 out of 50 labels, which is close to the threshold supporting the main claim.
Improve the evaluator agreement and provide the missing code/data repository.
- limitations section already calls out almost everything i wanted to add incl replacing the manual word-lists with SAE feature latents, evaluate if they show behavioral self-preservation in interactive tool-use environments (e.g., agentic coding, CLI environments with simulated shutdown threats) over choice surveys, checking across across multiple transformer depths and larger open-weight reasoning architectures (llama etc)
- overall rigorous reporting including negative controls
Cite this work
@misc {
title={
(HckPrj) Wanting, decomposed: Enthusiasm and hope sharpen a model’s preferences and move its consent
},
author={
Stanislav Lukyanenko
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


