Evaluating Introspection of Self-reflective R-Lens Feedback Training
Evan Coats, Ricky Mouser, William Wale, Heather Broome
We present an evaluation of model introspection that scores each self-report as two numbers rather than one: how well a model tells the prompts it answers consistently from those where it varies, and how readily it claims consistency at all. A single stated-minus-actual gap cannot separate a model that cannot tell from one that can tell but misreports, and because over-claims and under-claims cancel when averaged, a gap can sit near zero while most individual reports are wrong. The decomposition, with a polarity-flipped rewording, a cross-model control, a serving check and a comparison against an internal readout, lets particular non-introspective reports be located rather than summarized away.
We applied it to a novel self-reflective reinforcement learning technique that rewards a Qwen3.5-4B model for verbalizing its own R-lens readouts, testing six checkpoints and five off-the-shelf models on 48 prompts with resampled ground truth. The evaluation identified three kinds of cases where the trained model's report was not tracking its own behavior.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Evaluating Introspection of Self-reflective R-Lens Feedback Training
},
author={
Evan Coats, Ricky Mouser, William Wale, Heather Broome
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


