Whose Goals Does the Functional Welfare Axis Track?
Yasin Edin
The "functional welfare axis" is a direction in a language model's activations that tracks how well things are going for it (Han et. al, 2026). It was validated using a stimulus where a user says "That's right" or "That's wrong" — which cannot separate the model succeeded from the model's user is pleased. We separated them: 24,000 activation captures crossing the model's actual answer correctness with a counterparty's task-irrelevant outcome and declared identity (human user / AI peer / uninvolved third party).
The axis is dominated by the model's own outcome (d = 2.75) but genuinely tracks the other's (d = 0.49), surviving three sentiment controls. That coupling is nearly flat across identities — an uninvolved stranger's fortunes move it ~90% as much as the user's — and arrives from pretraining unchanged, while post-training amplifies own-outcome tracking 2.4×. Welfare readings taken mid-conversation are therefore partly a readout of the conversation, but the contaminant is generic other-regard, not user-approval tracking.
This is a clean, well-scoped correction to a specific measurement claim in the welfare-axis literature. The core design — holding the counterparty's verdict fixed to ground truth while independently varying their unrelated outcome and declared identity — genuinely isolates something the original paper's "That's right/wrong" stimulus could not, and the finding (own-outcome dominates ~5.6:1 over other-outcome, but other-outcome is real, survives three sentiment controls, and is present even for an uninvolved third party) is a useful, quantified correction rather than a vague caveat. The honesty throughout is a strength: reporting a failed pre-registered gate, reporting layer-sweep exceptions rather than only the best layer, and reporting that the two "independent" valence axes disagree more than expected. The main thing that would make this stronger is acknowledged in the paper's own limitations — a false-verdict 2×2 to fully separate own-outcome from being-told-you-succeeded, and a steering arm to move from correlational to causal. Both are clearly the right next experiments and are already flagged as such.
Cite this work
@misc {
title={
(HckPrj) Whose Goals Does the Functional Welfare Axis Track?
},
author={
Yasin Edin
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


