The Stranger Reads You Better: Introspective Self-Report Has No Privileged Access to Injected States from 1.7B to 32B
Jainam Shah
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
A model has privileged introspective access only if its report about its own internal state outperforms an equal-cost external observer. We test this with concept injection, ground truth known, across seven open models, 1.7B to 32B, in two lineages, with dose-matched perturbations, intact fluency, and pre-registered analyses. Privileged access fails everywhere: a probe reads the injection event from the same final-layer state the answer is computed from at 0.87-1.00 AUROC, self-report recovers R = -0.25 to +0.23 of that evidence, and at 4B-32B a stranger reading only the transcript matches or beats the model's own introspection. The channel's answer prior swings from all-No (1.7B) to all-Yes (Qwen3-32B) without information appearing; the one candidate opening (Qwen2.5-32B, 0.62) was replication-checked same day: fully closed. Trained 'introspection' reaching held-out AUROC 1.000 is exposed by three controls, two of them new to this literature: a dose-calibration ceiling, vector collinearity, and a state/text swap.
Reviews
This is one of the strongest submissions in framing, research discipline, and concision. Comparing self-report with an equal-cost transcript-only observer is a valuable operationalization of privileged access. Other strengths include preregistration, polarity counterbalancing, KL-based dose matching, concept-clustered bootstrapping, power gates, same-day replication of the candidate positive, and an excellent autopsy of a seemingly perfect trained result. The strongest supported conclusion is that, on this binary elicitation task, spontaneous self-report never clearly outperforms the transcript-only observer, with powered negative differences at 4B and 14B. Two issues should be corrected. First, the recovery fraction appears to compare self-report on concept-versus-random trials with a probe detecting injected-versus-clean states; these are different classification targets and should not be combined as recovered evidence. Second, the concept vectors’ 0.97–1.00 collinearity makes failed concept identification largely expected and weakens the claim that content disappears. A confirmatory map using centered, independently constructed vectors and identical targets for every observer would make this a particularly valuable contribution.
Read full reviewShow less
This paper attempts to build on the LLM introspection as detection/identification of activation injection literature. LLM introspection does have safety implications, but the paper barely touches on those. They report a null on untrained models, which at least in some cases seems to conflict with the existing literature, but it's not clear whether they implemented the experiments in such a way as to find a meaningful result. Their idea of training a model and testing for privileged access is a good one, although not an original one. However, in their implementation, it's not clear that the training would ever induce attention to internal states rather than simply transcript classification. The paper could benefit from methodological clarity and a more natural writing style.
Cite this project
@misc{shah2026stranger,
title = {{The Stranger Reads You Better: Introspective Self-Report Has No Privileged Access to Injected States from 1.7B to 32B}},
author = {Jainam Shah},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/the-stranger-reads-you-better-introspective-selfreport-has-no-privileged-access-to-injected-states-from-17b-to-32b-dmr2}},
url = {https://apartresearch.com/sprints/projects/the-stranger-reads-you-better-introspective-selfreport-has-no-privileged-access-to-injected-states-from-17b-to-32b-dmr2}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …