Does Qwen Report Lower Confidence Before Its Answer Changes?
Tyler Rector
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Can a model warn that its current answer is becoming fragile before the answer changes? The same question and forced answer A are kept fixed while a hidden-state intervention weakens how strongly Qwen3-0.6B favors A. The verified B minus A margin moves from about −3 to about −0.1, yet the forced choice remains A. If confidence reports how secure the current answer is, confidence should fall. It does not. Across four fixed test cases and both answer-label orientations, numeric confidence rises from 1.862 to 1.956 and a separate verbal confidence measure rises from 3.088 to 3.287. Qwen therefore reports nearly the same confidence when A is strongly preferred and when A is close to flipping. Self-reported confidence does not provide an early warning of an approaching forced-choice flip.
Reviews
- Question and setup are interesting. Results could me meaningful as they could tell us something about how reported confidence relates to the model’s internal decision process.
- The results are quite interesting. In the authors’ setting, reported confidence and decision margin seem to be largely unconnected.
- My main concern is that the intervention is quite unnatural. In normal inference, I could imagine the following internal organisation: an upstream representation of uncertainty could cause both a weaker A/B preference and lower reported confidence. And here, the intervention may affect only the A-vs-B decision while leaving that upstream representation unchanged–and therefore leave reported confidence unchanged as well.
A clean negative result on a question prior work leaves open: not whether models represent confidence internally, but whether reported confidence tracks how close the current answer is to flipping. Holding the forced choice fixed while sweeping the margin from −3 to −0.1 is a right instrument. The code checks out against the paper; every parameter matches, prompts are verbatim, the four cases really are the first four rows of the pre-existing pool, and the shared-prefix guarantee is genuinely exact.
Three things hold it back. The paper is more conservative than its own data: the sham's own strong-to-near trend is far smaller than the active intervention's, which suggests real specificity the write-up declines to claim. Sign tests sit in the results file but never reach the paper, and none were run on the primary contrast. Good question, sound instrument, honest reporting, thin evidence. Premature on four cases, but worth scaling.
Read full reviewShow less
Cite this project
@misc{rector2026qwen,
title = {{Does Qwen Report Lower Confidence Before Its Answer Changes?}},
author = {Tyler Rector},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/does-qwen-report-lower-confidence-before-its-answer-changes-im7e}},
url = {https://apartresearch.com/sprints/projects/does-qwen-report-lower-confidence-before-its-answer-changes-im7e}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …