The Introspection Gap: A Trained Probe Recovers What Self- Report Misses
Omanshu Thapliyal
AI safety research usually treats a language model’s self-reports about its own processing as evidence about what is happening inside it. Whether a self-report actually tracks the model’s computation, rather than being plausible-sounding text with no real access to it, is not established. We test this with two ground-truth paradigms: activation injection, which perturbs internal representations directly and asks whether the model notices, and context injection, which places a fabricated fact in conversation and asks whether an answer depended on it. Activation-injection self-report is a comprehensive null across every architecture and training regime tested, including a positive-control sweep to six times the tested injection range, corroborated by a non-linguistic detection method with no dependence on language output at all. Context-injection self-report shows a real positive signal, but a trained linear probe on the same internal state detects the ground truth far more reliably than self-report does (AUC 0.86–0.95 versus accuracy never exceeding 62.5%, across four models). Decomposing introspective-question wording into six structural axes, three (length, formality, reflective framing) significantly affect reliability, while the axis an earlier check had credited most does not replicate at scale. Self-report reliability therefore depends on which paradigm and question is used, not so much on architecture or scale, and a model’s internal state is often more informative than the model itself.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) The Introspection Gap: A Trained Probe Recovers What Self- Report Misses
},
author={
Omanshu Thapliyal
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


