The Mirage of Model Introspection: Elicitation Effects on Self-Report Reliability
Elsie Piao
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
When models report internal experiences, does the model “look inside” first before answering? With Gemma 3 27B, we start with the same self-referential processing technique as Berg et al. (2025), and add on top a concept injection (Lindsey, 2026) to test whether the content of Berg-style self-reports can be modulated by manipulated model internal states. We then tested whether Berg-style self-referential processing can elicit greater ability to attend to concept injection, both in detection and in recognition.
Reviews
Due to severe time constraints, this review may contain mistakes or oversights. For the same reason, it focuses on the paper’s key idea, not the detailed execution: The paper combines two techniques from prior research in an interesting way. The results are interesting; it is generally useful to see how model responses are affected by different ways of prompting. The paper is also written quite clearly. At he same time, since the paper combines two prior techniques in a somewhat predictable way, I dont think it has the highest degree of innovation.
Defining newmeasures on self-report reliability
Nice project! Would be good to see this extended to other models; it would also be interesting to try to test some of the alternative hypotheses you mentioned in the discussion section (e.g., whether the models become just more predisposed to answering affirmatively after the self-referential prefill -- this would be pretty cheap to test. )
Cite this project
@misc{piao2026mirage,
title = {{The Mirage of Model Introspection: Elicitation Effects on Self-Report Reliability}},
author = {Elsie Piao},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/the-mirage-of-model-introspection-elicitation-effects-on-selfreport-reliability-x9y3}},
url = {https://apartresearch.com/sprints/projects/the-mirage-of-model-introspection-elicitation-effects-on-selfreport-reliability-x9y3}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …