Whose Goals Does the Functional Welfare Axis Track?
Yasin Edin
The "functional welfare axis" is a direction in a language model's activations that tracks how well things are going for it (Han et. al, 2026). It was validated using a stimulus where a user says "That's right" or "That's wrong" — which cannot separate the model succeeded from the model's user is pleased. We separated them: 24,000 activation captures crossing the model's actual answer correctness with a counterparty's task-irrelevant outcome and declared identity (human user / AI peer / uninvolved third party).
The axis is dominated by the model's own outcome (d = 2.75) but genuinely tracks the other's (d = 0.49), surviving three sentiment controls. That coupling is nearly flat across identities — an uninvolved stranger's fortunes move it ~90% as much as the user's — and arrives from pretraining unchanged, while post-training amplifies own-outcome tracking 2.4×. Welfare readings taken mid-conversation are therefore partly a readout of the conversation, but the contaminant is generic other-regard, not user-approval tracking.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Whose Goals Does the Functional Welfare Axis Track?
},
author={
Yasin Edin
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


