Whose Goals Does the Functional Welfare Axis Track?
Yasin Edin · Team Whose Welfare
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
The "functional welfare axis" is a direction in a language model's activations that tracks how well things are going for it (Han et. al, 2026). It was validated using a stimulus where a user says "That's right" or "That's wrong" — which cannot separate the model succeeded from the model's user is pleased. We separated them: 24,000 activation captures crossing the model's actual answer correctness with a counterparty's task-irrelevant outcome and declared identity (human user / AI peer / uninvolved third party).
The axis is dominated by the model's own outcome (d = 2.75) but genuinely tracks the other's (d = 0.49), surviving three sentiment controls. That coupling is nearly flat across identities — an uninvolved stranger's fortunes move it ~90% as much as the user's — and arrives from pretraining unchanged, while post-training amplifies own-outcome tracking 2.4×. Welfare readings taken mid-conversation are therefore partly a readout of the conversation, but the contaminant is generic other-regard, not user-approval tracking.
Reviews
This is a clean, well-scoped correction to a specific measurement claim in the welfare-axis literature. The core design — holding the counterparty's verdict fixed to ground truth while independently varying their unrelated outcome and declared identity — genuinely isolates something the original paper's "That's right/wrong" stimulus could not, and the finding (own-outcome dominates ~5.6:1 over other-outcome, but other-outcome is real, survives three sentiment controls, and is present even for an uninvolved third party) is a useful, quantified correction rather than a vague caveat. The honesty throughout is a strength: reporting a failed pre-registered gate, reporting layer-sweep exceptions rather than only the best layer, and reporting that the two "independent" valence axes disagree more than expected. The main thing that would make this stronger is acknowledged in the paper's own limitations — a false-verdict 2×2 to fully separate own-outcome from being-told-you-succeeded, and a steering arm to move from correlational to causal. Both are clearly the right next experiments and are already flagged as such.
Read full reviewShow less
Cite this project
@misc{edin2026whose,
title = {{Whose Goals Does the Functional Welfare Axis Track?}},
author = {Yasin Edin},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/whose-goals-does-the-functional-welfare-axis-track-6ae5}},
url = {https://apartresearch.com/sprints/projects/whose-goals-does-the-functional-welfare-axis-track-6ae5}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …