Skip to content
Sprint projectAug 16, 2026BENGALURU

Is Valence in the Global Workspace?

Ujjwal Kukreti · Team Tachyon

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

LLM self-reports could support monitoring during conversations, but a report may reflect the prompt rather than the model’s internal activation. We investigate whether a valence-related activation direction causally influences self-reports and behavior. Across four open-weight models, we extract and validate a direction that separates positive from negative scenarios, then intervene on it using activation steering. We test self-report, generated-response tone, and refusal-like CONTINUE/EXIT decisions. The direction is strongly decodable in every model, and steering consistently changes response tone. However, natural correlation with self-report does not predict causal sensitivity: Phi-3-mini shows the highest correlation but follows steering in only 5.6% of conflict trials, while SmolLM2 follows steering in 62.0%. These results show that self-reports should be validated through intervention, not correlation alone.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The submission asks whether a numerical valence self-report from a small open-weight language model is a genuine readout of an internal valence representation or a reconstruction from the prompt, and it answers with a conflict design that holds a valence-laden scenario fixed while steering a decoded valence direction in the opposite sign, reporting that the models whose reports correlate most strongly with the decoded state are the ones whose reports follow the steering least. The design is the right instrument for that question, the supporting controls for random directions, affect-free wording, system personas and steering dose go beyond what a 3-day research sprint usually delivers, and the reported proportions and their intervals recompute correctly from the stated denominators. The most valuable next step would be to plot the follow-steering rate against parameter count, because the four models are ordered by size in the same way they are ordered by report correlation, so the reported inversion and the more mundane reading that larger models resist a fixed-magnitude perturbation cannot currently be told apart. A within-family comparison at steering magnitudes calibrated to be equipotent across models would begin to separate the two accounts. A second step is a positive control demonstrating that some activation intervention can move the numerical report in the model whose report never moves, since without one the central negative result cannot be distinguished from an intervention that was simply too weak at the single layer and magnitude chosen.

    Read full reviewShow less

Cite this project

@misc{kukreti2026valence,
  title = {{Is Valence in the Global Workspace?}},
  author = {Ujjwal Kukreti},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/is-valence-in-the-global-workspace-2vsj}},
  url = {https://apartresearch.com/sprints/projects/is-valence-in-the-global-workspace-2vsj}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026