Causal confidence steering supresses metacognitive error detection
Trevor Harrison
We causally raised an internal confidence signal in Qwen2.5-7B-Instruct while holding the question and answer fixed, and found it made the model less likely to flag a plainly false answer as inconsistent (−.41) or incorrect (−.50). Confidence appears to govern the model's self-monitoring rather than being monitored by it, and we find no behavioral evidence that the model notices when its confidence has been artificially inflated.
The paper’s results support the title. Adding a learned confidence-related activation direction into Qwen2.5-7B-Instruct changes both reported confidence and reports of inconsistency/error. In plain English, when we artificially raise an internal state associated with confidence while leaving the question and answer unchanged, the model becomes less likely to recognise that the answer is wrong. Separating answer quality from confidence in this way is a helpful incremental contribution to the key question of “when models know that they don’t know”.
The results may or may not be a specific error-detection metacognition mechanism (vs level of ‘verbal commitment’ in response to RLHF priorities) – some of the paper’s phrasing is over-confident in this regard, but the overall limitations and caveats are well described. It’s worth noting that the experimenter activation appears quite large in scale, so it has taken a significant intervention to shape the model’s reports.
The results are worth replicating on other models with different parameter choices (there are significant researcher degrees of freedom in the method), alongside extensions such as norm-matched placebo interventions and identifying ways for the model to generate its own mistake for evaluation (as opposed to experimenter imposed mistakes). It would also be helpful to translate intervention strengths (alpha parameter) and impacts on reports into intuitive language that are contextualised against ordinary model operation. Well known evidence elsewhere in the field that some models can detect injections would also be worth exploring (e.g. is it detected in these cases and, when it is detected, does that mediate the target relationship).
It is worth noting the paper assumes significant knowledge of prior techniques and the text reads as quite compressed, so it may hard to follow in full for a general AI safety audience, although readers can follow the citations to gain the additional context.
Cite this work
@misc {
title={
(HckPrj) Causal confidence steering supresses metacognitive error detection
},
author={
Trevor Harrison
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


