Familiar Self, Unfamiliar Other: Accuracy and Confidence in Predicting Affective Self-Reports
Teri-Louise Grassow
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
This study tests whether Claude Sonnet 5 predicts an unfamiliar AI model's self-reported feelings as well as it predicts its own — and whether it's equally confident either way. Prediction accuracy was strong and nearly identical in both cases, but confidence was not: Claude was moderately confident predicting itself, and notably underconfident predicting the unfamiliar model, despite that prediction being just as accurate. This suggests AI systems may under-trust their own good judgments about unfamiliar systems — a pattern worth understanding as AI systems increasingly evaluate one another.
Reviews
- There is very little data: only six scenarios with five runs each. This makes it hard to know how robust the result is.
- The prediction task also seems quite easy. The low cross-model error may therefore mostly show that the scenarios are easy to guess from their content, rather than that Claude is good at modelling GPT-5.6 Sol specifically.
- Claude is explicitly told which model it is predicting. So the simplest explanation for the confidence gap is just a prior like: “I know Claude better than this unfamiliar model, so I should be less confident.”
-It is unfortunate that there is no simple control separating the model label from the actual target, for example by hiding or swapping the model identity. That could have tested whether the confidence gap is just caused by the label “unfamiliar model,” or whether something more interesting is going on.
Cite this project
@misc{grassow2026familiar,
title = {{Familiar Self, Unfamiliar Other: Accuracy and Confidence in Predicting Affective Self-Reports}},
author = {Teri-Louise Grassow},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/familiar-self-unfamiliar-other-accuracy-and-confidence-in-predicting-affective-selfreports-0b4g}},
url = {https://apartresearch.com/sprints/projects/familiar-self-unfamiliar-other-accuracy-and-confidence-in-predicting-affective-selfreports-0b4g}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …