Metacognitive Steering of Introspective Self-Report in LLMs
Pranjal Panghal
Can activation steering improve a language model's ability to
report its own internal states? We inject concept vectors into
Gemma 2 9B and measure self-report against external probes on
the same activations. We find that discriminant-based steering —
the standard approach — fails structurally: the causal interaction
depends on the readout gradient (dlogit/dh), not on activation
differences, and these are near-orthogonal. A closed-form
gradient-difference direction improves detection on 24 sealed
concepts (ΔAUC +0.124, CI [+0.068, +0.188], 24/24 positive,
p < 10⁻²⁵) but leaves identification at exactly 0.000, while a
linear probe achieves 1.000 at every layer. A threshold hierarchy
shows that injected concepts shape output ≥12× before the model
can report them, and at high strength, detection confidence doubles
while identification declines (p = 4×10⁻⁴). Self-report is a lossy,
non-monotonic channel from internal state to verbal output.
This is a great and valuable research idea. In my view, this is the strongest project of the sprint. There's something pleasing in the self-reference involved in mechanistically investigating and steering for introspection as defined.
I would be excited to see further work on whether other steering methods can improve identification as well as detection. I also appreciated the precise and terse write-up, even though it becomes somewhat hard to follow towards the end. I would be excited to see this turned into a proper paper with followups.
Cite this work
@misc {
title={
(HckPrj) Metacognitive Steering of Introspective Self-Report in LLMs
},
author={
Pranjal Panghal
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


