Skip to content
Sprint projectAug 17, 2026Edmonton

The Stranger Reads You Better: Introspective Self-Report Has No Privileged Access to Injected States from 1.7B to 32B

Jainam Shah

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: The Stranger Reads You Better: Introspective Self-Report Has No Privileged Access to Injected States from 1.7B to 32B

Code (opens in new tab)
Share

A model has privileged introspective access only if its report about its own internal state outperforms an equal-cost external observer. We test this with concept injection, ground truth known, across seven open models, 1.7B to 32B, in two lineages, with dose-matched perturbations, intact fluency, and pre-registered analyses. Privileged access fails everywhere: a probe reads the injection event from the same final-layer state the answer is computed from at 0.87-1.00 AUROC, self-report recovers R = -0.25 to +0.23 of that evidence, and at 4B-32B a stranger reading only the transcript matches or beats the model's own introspection. The channel's answer prior swings from all-No (1.7B) to all-Yes (Qwen3-32B) without information appearing; the one candidate opening (Qwen2.5-32B, 0.62) was replication-checked same day: fully closed. Trained 'introspection' reaching held-out AUROC 1.000 is exposed by three controls, two of them new to this literature: a dose-calibration ceiling, vector collinearity, and a state/text swap.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This is one of the strongest submissions in framing, research discipline, and concision. Comparing self-report with an equal-cost transcript-only observer is a valuable operationalization of privileged access. Other strengths include preregistration, polarity counterbalancing, KL-based dose matching, concept-clustered bootstrapping, power gates, same-day replication of the candidate positive, and an excellent autopsy of a seemingly perfect trained result. The strongest supported conclusion is that, on this binary elicitation task, spontaneous self-report never clearly outperforms the transcript-only observer, with powered negative differences at 4B and 14B. Two issues should be corrected. First, the recovery fraction appears to compare self-report on concept-versus-random trials with a probe detecting injected-versus-clean states; these are different classification targets and should not be combined as recovered evidence. Second, the concept vectors’ 0.97–1.00 collinearity makes failed concept identification largely expected and weakens the claim that content disappears. A confirmatory map using centered, independently constructed vectors and identical targets for every observer would make this a particularly valuable contribution.

    Read full reviewShow less
  2. This paper attempts to build on the LLM introspection as detection/identification of activation injection literature. LLM introspection does have safety implications, but the paper barely touches on those. They report a null on untrained models, which at least in some cases seems to conflict with the existing literature, but it's not clear whether they implemented the experiments in such a way as to find a meaningful result. Their idea of training a model and testing for privileged access is a good one, although not an original one. However, in their implementation, it's not clear that the training would ever induce attention to internal states rather than simply transcript classification. The paper could benefit from methodological clarity and a more natural writing style.

Cite this project

@misc{shah2026stranger,
  title = {{The Stranger Reads You Better: Introspective Self-Report Has No Privileged Access to Injected States from 1.7B to 32B}},
  author = {Jainam Shah},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/the-stranger-reads-you-better-introspective-selfreport-has-no-privileged-access-to-injected-states-from-17b-to-32b-dmr2}},
  url = {https://apartresearch.com/sprints/projects/the-stranger-reads-you-better-introspective-selfreport-has-no-privileged-access-to-injected-states-from-17b-to-32b-dmr2}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026