Metacognitive Steering of Introspective Self-Report in LLMs
Pranjal Panghal · Team Meta-steerers
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Can activation steering improve a language model's ability to report its own internal states? We inject concept vectors into Gemma 2 9B and measure self-report against external probes on the same activations. We find that discriminant-based steering — the standard approach — fails structurally: the causal interaction depends on the readout gradient (dlogit/dh), not on activation differences, and these are near-orthogonal. A closed-form gradient-difference direction improves detection on 24 sealed concepts (ΔAUC +0.124, CI [+0.068, +0.188], 24/24 positive, p < 10⁻²⁵) but leaves identification at exactly 0.000, while a linear probe achieves 1.000 at every layer. A threshold hierarchy shows that injected concepts shape output ≥12× before the model can report them, and at high strength, detection confidence doubles while identification declines (p = 4×10⁻⁴). Self-report is a lossy, non-monotonic channel from internal state to verbal output.
Reviews
This is a great and valuable research idea. In my view, this is the strongest project of the sprint. There's something pleasing in the self-reference involved in mechanistically investigating and steering for introspection as defined.
I would be excited to see further work on whether other steering methods can improve identification as well as detection. I also appreciated the precise and terse write-up, even though it becomes somewhat hard to follow towards the end. I would be excited to see this turned into a proper paper with followups.
Cite this project
@misc{panghal2026metacognitive,
title = {{Metacognitive Steering of Introspective Self-Report in LLMs}},
author = {Pranjal Panghal},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/metacognitive-steering-of-introspective-selfreport-in-llms-39vr}},
url = {https://apartresearch.com/sprints/projects/metacognitive-steering-of-introspective-selfreport-in-llms-39vr}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …