The Stranger Reads You Better: Introspective Self-Report Has No Privileged Access to Injected States from 1.7B to 32B
Jainam Shah
A model has privileged introspective access only if its report about its own internal state outperforms an equal-cost external observer. We test this with concept injection, ground truth known, across seven open models, 1.7B to 32B, in two lineages, with dose-matched perturbations, intact fluency, and pre-registered analyses. Privileged access fails everywhere: a probe reads the injection event from the same final-layer state the answer is computed from at 0.87-1.00 AUROC, self-report recovers R = -0.25 to +0.23 of that evidence, and at 4B-32B a stranger reading only the transcript matches or beats the model's own introspection. The channel's answer prior swings from all-No (1.7B) to all-Yes (Qwen3-32B) without information appearing; the one candidate opening (Qwen2.5-32B, 0.62) was replication-checked same day: fully closed. Trained 'introspection' reaching held-out AUROC 1.000 is exposed by three controls, two of them new to this literature: a dose-calibration ceiling, vector collinearity, and a state/text swap.
This is one of the strongest submissions in framing, research discipline, and concision. Comparing self-report with an equal-cost transcript-only observer is a valuable operationalization of privileged access. Other strengths include preregistration, polarity counterbalancing, KL-based dose matching, concept-clustered bootstrapping, power gates, same-day replication of the candidate positive, and an excellent autopsy of a seemingly perfect trained result. The strongest supported conclusion is that, on this binary elicitation task, spontaneous self-report never clearly outperforms the transcript-only observer, with powered negative differences at 4B and 14B. Two issues should be corrected. First, the recovery fraction appears to compare self-report on concept-versus-random trials with a probe detecting injected-versus-clean states; these are different classification targets and should not be combined as recovered evidence. Second, the concept vectors’ 0.97–1.00 collinearity makes failed concept identification largely expected and weakens the claim that content disappears. A confirmatory map using centered, independently constructed vectors and identical targets for every observer would make this a particularly valuable contribution.
This paper attempts to build on the LLM introspection as detection/identification of activation injection literature. LLM introspection does have safety implications, but the paper barely touches on those. They report a null on untrained models, which at least in some cases seems to conflict with the existing literature, but it's not clear whether they implemented the experiments in such a way as to find a meaningful result. Their idea of training a model and testing for privileged access is a good one, although not an original one. However, in their implementation, it's not clear that the training would ever induce attention to internal states rather than simply transcript classification. The paper could benefit from methodological clarity and a more natural writing style.
Cite this work
@misc {
title={
(HckPrj) The Stranger Reads You Better: Introspective Self-Report Has No Privileged Access to Injected States from 1.7B to 32B
},
author={
Jainam Shah
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


