Having a State Is Not Knowing It
Amrit Gopinath, Raghul Sugumar · Team Layer 8 Legends
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
"Having a State Is Not Knowing It" treats introspection as a hierarchy of falsifiable capabilities, not a binary trait. Replicating concept-injection in Llama-3.2-3B-Instruct shows injected concepts causally steer behavior and leave an attention trace even when verbal self-report fails. A controlled Qwen decision-state benchmark, with ground truth from measured logit-margin shifts, shows native self-report is weak (F1 0.24) but trainable (F1 0.52) — though external probes decode the same state perfectly, ruling out privileged access. A counterfactual protocol shows the model predicts intervention direction (F1 0.70) but not magnitude: quantitative self-modeling fails cleanly (negative R²).
Reviews
The five-gate hierarchy is highly useful, and other researchers should adopt it. The authors built trust by avoiding easy-to-cheat testing methods. However, the core finding has a small margin of error and flips with new data mixtures. To strengthen the paper, the authors should add repeated tests, confidence intervals, and open-source data. Finally, a clear flowchart and simpler formatting would make it easier to read
This work asks a very interesting question: whether language models can predict how hypothetical interventions to their own internal states would affect their behavior. The experimental design is thoughtful, though a schematic figure of the intervention, prediction, and scoring setup would make it much easier to follow.
One concern is that poor introspective performance after an intervention may not necessarily indicate limited introspective ability, since the steering itself could disrupt the computation needed for introspection or accurate reporting. A related concern also appears in the cited work. It would be interesting to test whether allowing the model additional computation or reasoning before reporting mitigates this effect.
The limited OOD generalization also leaves open whether the observed failures reflect a fundamental limitation or insufficient training, data, or model capacity; scaling these factors would help distinguish the two. Finally, the weak prediction of quantitative effect size may partly reflect the difficulty of mapping internal changes onto an abstract continuous magnitude. A simpler ordinal scale (e.g., no effect / small / medium / large effect) might provide a more natural test of whether the model can estimate intervention strength.
Read full reviewShow less
Cite this project
@misc{gopinath2026having,
title = {{Having a State Is Not Knowing It}},
author = {Amrit Gopinath and Raghul Sugumar},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/having-a-state-is-not-knowing-it-0zbf}},
url = {https://apartresearch.com/sprints/projects/having-a-state-is-not-knowing-it-0zbf}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …