Skip to content
Sprint projectAug 17, 2026Bangalore
1st place

Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models

Vishwa Kumaresh

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models

Code (opens in new tab)
Share

Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is unresolved. We introduce OWL, a factorial SELF × OTHER × surface-affect benchmark with held-out scenario families and rendering controls. Using linear probes, factorial representation decomposition, and natural counterfactual activation interchange in a learned attribution subspace, we evaluate predictive and causal attribution of functional valence. Linear probes decode the active assistant's functional outcome with test AUROC 0.859 versus 0.601 for a represented other and survive every rendering control (0.848–0.999), yet causal interchange in the learned S subspace recovers only 0.18 of the natural counterfactual effect on self-directed choice (full-vector patch 0.42, random ≈ 0). Readability without causal efficacy implies that internal valence probes require intervention-based validation.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Motivation for the project is clear and effectively argued. Probes for internal welfare states don't distinguish between representations of the model's own welfare states and others welfare states. Authors show that the model can distinguish ownership but that the model's "own" welfare states don't disproportionately cause its own actions. This is very useful work, though more investigation is needed into the causal relationships would be between model welfare states (if any exist) and model behavior, given assistant training, safety training, and other protocols that may prevent models from acting on their own states. It would have been good to see more detailed presentation of the OWL benchmark used.

  2. This is an unusually strong sprint submission that addresses an important construct-validity problem for AI welfare measurement: distinguishing whether a valenced representation belongs to the active assistant or merely to another represented entity. I particularly appreciated the factorial SELF × OTHER × affect design, held-out scenario families and unseen renderings, lexical baselines, natural-counterfactual activation interchange, full-vector and random controls, and the explicit separation between predictive decodability and causal use.

    The "readable but not causal" result is compelling: SELF outcome is strongly decodable across held-out conditions, yet intervention in the learned SELF subspace recovers only a small fraction of the natural behavioral effect. The fact that the full-vector patch recovers more and random controls are near null makes the negative causal result substantially more informative than probe accuracy alone.

    I would nevertheless narrow the mechanistic interpretation slightly. Failure of the rank-1 SELF subspace demonstrates that this linear direction is not sufficient for the downstream behavior under the stated intervention, but it does not uniquely establish distributed/nonlinear owner binding. The full-vector intervention itself recovers only part of the natural counterfactual effect, so information outside the selected anchor/layer or other violations of the stated identification assumptions remain possible. Multi-layer/token patching and learned distributed subspaces such as Boundless DAS would be particularly useful follow-ups.

    Cross-family replication would also materially strengthen the result, as the current study is limited to Qwen3-8B. The ownership-conditioned monitoring result is promising but appropriately reported as nonsignificant; expanding the distress hard-negative set would help determine whether the apparent false-positive reduction generalizes.

    Overall, this is a rigorous and valuable contribution, especially because it demonstrates why strong probe performance should not automatically be interpreted as evidence that the decoded feature is the causal representation used by the model.

    Read full reviewShow less

Cite this project

@misc{kumaresh2026readable,
  title = {{Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models}},
  author = {Vishwa Kumaresh},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/readable-but-not-causal-limits-of-selfattributed-welfare-representations-in-language-models-ehin}},
  url = {https://apartresearch.com/sprints/projects/readable-but-not-causal-limits-of-selfattributed-welfare-representations-in-language-models-ehin}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026