Skip to content
Sprint projectAug 17, 2026Tokyo, Japan

Is "Confidence" the Right Word for Answer Repeat Probability?

Koshiro Aoki · Team bluetree

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Is "Confidence" the Right Word for Answer Repeat Probability?

Code (opens in new tab)
Share

Natural-language confidence reports depend on how they are elicited. We therefore test whether a single learned word can reliably report a frozen model’s answer repeat probability. For each multiple-choice question, the target is the probability that the model repeats its most likely valid answer. We learn one input embedding in Llama-3.1-8B and map the model’s 0–10 report to this target. Under a carrier shift, eight initializations yield Spearman correlations from 0.359 to 0.586. The reporter selected on the development set does not significantly outperform fixed confidence or consistency. A parameter-matched generic soft vector has lower mean squared error on both carriers, and readable replacements do not preserve the learned behavior. A learned vector can therefore adapt a particular reporting context, but these results do not establish a stable word-level interface or privileged introspection.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This work is motivated by an intriguing idea - that neologisms may allow us to more directly map between human intentions and the mental language of LLMs - which could have implications for model understanding and thereby AI safety to the extent it generalizes. It probes the special case of a model's ability to report the probability that it will repeat the same answer to a multiple choice question when asked again, and attempts to train a neologism for that. My assessment of the informativeness of the null result they report is constrained by my doubts about whether that is a viable target (it's not clear to me that the model would have that concept). However, the work is technically sound, uses appropriate baselines and controls, and overall offers a valuable methodological contribution. It could benefit from a larger set of carriers, diversity of training data, and attention to clarity of writing.

    Read full reviewShow less
  2. - The abstract is very difficult to follow. Terms such as "answer repeat probability", "carrier shift", "eight initializations", and "parameter-matched generic soft vector" are introduced before the reader has an intuitive understanding of the experiment. I think the paper would benefit enormously from first explaining the basic idea in ordinary language.

    - The introduction could do more to explain why this question is important, particularly in the context of digital minds. The target studied here, answer repeat probability, can already be calculated directly from the model's logits. I can see the value of using this as a controlled test case for whether one can construct a stable learned lexical interface to a model property, but the paper could make much clearer why success on this task would ultimately matter for questions about introspection, welfare, or digital minds more generally.

    - It would have been very helpful to give one complete worked example of the experimental procedure. For example: show an actual multiple-choice question, the model's answer probabilities, the resulting answer-repeat target, the full reporting prompt containing the learned embedding, and the numerical report produced by the model. I found it a bit difficult to reconstruct the experiment from the current methods section.

    - But the result itself is nice: a learned embedding can adapt the model’s reports to the target in the training context. Unfortunately, this does not generalise. Which is fine!

    Read full reviewShow less

Cite this project

@misc{aoki2026confidence,
  title = {{Is "Confidence" the Right Word for Answer Repeat Probability?}},
  author = {Koshiro Aoki},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/is-confidence-the-right-word-for-answer-repeat-probability-ot69}},
  url = {https://apartresearch.com/sprints/projects/is-confidence-the-right-word-for-answer-repeat-probability-ot69}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026