Causal confidence steering supresses metacognitive error detection
Trevor Harrison · Team Modulated Confidence
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
We causally raised an internal confidence signal in Qwen2.5-7B-Instruct while holding the question and answer fixed, and found it made the model less likely to flag a plainly false answer as inconsistent (−.41) or incorrect (−.50). Confidence appears to govern the model's self-monitoring rather than being monitored by it, and we find no behavioral evidence that the model notices when its confidence has been artificially inflated.
Reviews
The paper’s results support the title. Adding a learned confidence-related activation direction into Qwen2.5-7B-Instruct changes both reported confidence and reports of inconsistency/error. In plain English, when we artificially raise an internal state associated with confidence while leaving the question and answer unchanged, the model becomes less likely to recognise that the answer is wrong. Separating answer quality from confidence in this way is a helpful incremental contribution to the key question of “when models know that they don’t know”.
The results may or may not be a specific error-detection metacognition mechanism (vs level of ‘verbal commitment’ in response to RLHF priorities) – some of the paper’s phrasing is over-confident in this regard, but the overall limitations and caveats are well described. It’s worth noting that the experimenter activation appears quite large in scale, so it has taken a significant intervention to shape the model’s reports.
The results are worth replicating on other models with different parameter choices (there are significant researcher degrees of freedom in the method), alongside extensions such as norm-matched placebo interventions and identifying ways for the model to generate its own mistake for evaluation (as opposed to experimenter imposed mistakes). It would also be helpful to translate intervention strengths (alpha parameter) and impacts on reports into intuitive language that are contextualised against ordinary model operation. Well known evidence elsewhere in the field that some models can detect injections would also be worth exploring (e.g. is it detected in these cases and, when it is detected, does that mediate the target relationship).
It is worth noting the paper assumes significant knowledge of prior techniques and the text reads as quite compressed, so it may hard to follow in full for a general AI safety audience, although readers can follow the citations to gain the additional context.
Read full reviewShow less
Cite this project
@misc{harrison2026causal,
title = {{Causal confidence steering supresses metacognitive error detection}},
author = {Trevor Harrison},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/causal-confidence-steering-supresses-metacognitive-error-detection-ybb4}},
url = {https://apartresearch.com/sprints/projects/causal-confidence-steering-supresses-metacognitive-error-detection-ybb4}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …