Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Vishwa Kumaresh
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is unresolved. We introduce OWL, a factorial SELF × OTHER × surface-affect benchmark with held-out scenario families and rendering controls. Using linear probes, factorial representation decomposition, and natural counterfactual activation interchange in a learned attribution subspace, we evaluate predictive and causal attribution of functional valence. Linear probes decode the active assistant's functional outcome with test AUROC 0.859 versus 0.601 for a represented other and survive every rendering control (0.848–0.999), yet causal interchange in the learned S subspace recovers only 0.18 of the natural counterfactual effect on self-directed choice (full-vector patch 0.42, random ≈ 0). Readability without causal efficacy implies that internal valence probes require intervention-based validation.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
},
author={
Vishwa Kumaresh
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


