The Mirage of Model Introspection: Elicitation Effects on Self-Report Reliability
Elsie Piao
When models report internal experiences, does the model “look inside” first before answering? With Gemma 3 27B, we start with the same self-referential processing technique as Berg et al. (2025), and add on top a concept injection (Lindsey, 2026) to test whether the content of Berg-style self-reports can be modulated by manipulated model internal states. We then tested whether Berg-style self-referential processing can elicit greater ability to attend to concept injection, both in detection and in recognition.
Due to severe time constraints, this review may contain mistakes or oversights. For the same reason, it focuses on the paper’s key idea, not the detailed execution: The paper combines two techniques from prior research in an interesting way. The results are interesting; it is generally useful to see how model responses are affected by different ways of prompting. The paper is also written quite clearly. At he same time, since the paper combines two prior techniques in a somewhat predictable way, I dont think it has the highest degree of innovation.
Defining newmeasures on self-report reliability
Nice project! Would be good to see this extended to other models; it would also be interesting to try to test some of the alternative hypotheses you mentioned in the discussion section (e.g., whether the models become just more predisposed to answering affirmatively after the self-referential prefill -- this would be pretty cheap to test. )
Cite this work
@misc {
title={
(HckPrj) The Mirage of Model Introspection: Elicitation Effects on Self-Report Reliability
},
author={
Elsie Piao
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


