Valence Lens: An Internal Valence Signal That Scales When Self-Report Does Not
Karan Singh, Shivansh Shukla · Team PROBE — Principal Representation & Objective Behavior Evaluation
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Can we detect an AI's internal "good/bad" state without asking it? Current welfare assessments rely on self-reports, which often reflect role-play, training pressure, or sycophancy rather than internal states. Prior work identifies the missing step: correlating responses with internal activations. ValenceLens supplies this independently and broadly.
ValenceLens is an open-weights probe reading a linear flourish-versus-distress (valence) direction from a model's residual stream via content-matched contexts. Across 12 instruct models (6 architectures, 8 organizations, 0.5B–7B), it separates valence with large effects (held-out Cohen's d 4.6–14). The direction is causal (steering beats 24 placebos, z 8–18, 11/12 models), arousal-orthogonal, irreducible to sentiment (~33% of variance), and generalizes to two human-labeled datasets (AUROC 0.85–0.89). It independently triangulates internal, verbal, and behavioral evidence, matching the field's methodological demands.
The central finding is an asymmetry: the internal signal is scale-invariant, while a model's ability to verbalize or act on its valence emerges only with capability. The probe therefore works exactly where self-report fails small models that cannot describe their own state, where behavioral audits are blindest. It runs on a 6 GB consumer laptop at $0 API cost: a cheap, independent welfare-audit primitive for third-party auditors, safety teams, and regulators.
Every prediction was pre-registered; we honestly report one overturned follow-up, a retracted claim, and a null introspection result. The signal is operationalized, without claiming about subjective experience.
Reviews
Investigating methodologies for measuring valence seems important. This work trains probes that classify valence. I felt methodological details were missing, such that I was not able to assess the validity of the results (e.g., what prompts are used for extracting the direction? what task is used for evaluation?).
Cite this project
@misc{singh2026valence,
title = {{Valence Lens: An Internal Valence Signal That Scales When Self-Report Does Not}},
author = {Karan Singh and Shivansh Shukla},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/valence-lens-an-internal-valence-signal-that-scales-when-selfreport-does-not-avrv}},
url = {https://apartresearch.com/sprints/projects/valence-lens-an-internal-valence-signal-that-scales-when-selfreport-does-not-avrv}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …