Where Self-Knowledge Fails. Models Predict Their Own Choices Well, but Misreport the Ones That Concern Themselves
Arpit Singh Gautam · Team ASG
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
AI-welfare claims rest on two untested and separable assumptions, that a model's stated evaluations match the choices it actually makes, and that it knows itself better than an outside observer does. We test both using three independent elicitations over identical material, namely forced pairwise choice, one-at-a-time cardinal rating, and predicted choice, and add a concept-injection benchmark with ground truth. Stated and revealed preferences agree well overall, at 0.872 and 0.828, but diverge sharply on outcomes concerning the model itself, the lowest-agreeing substantive category in both models at 0.643 and 0.548. The standard cross-model test of privileged access is confounded, because a noisier external predictor scores lower at predicting any target. Under a within-model contrast holding instrument quality fixed, Qwen2.5-7B retains a small advantage of 0.031 and Mistral-7B does not. On injection, the two highest raw detection rates belong to models that report an injected concept more than half the time when nothing is injected.

Reviews
You've taken the common assumption (that models have privileged access to their own preferences) and run it through a proper experimental wringer. The big headline is that most of what looked like self-knowledge is actually just the model being good at predicting what any assistant would do. Your within-model contrast is the methodological innovation and the injection results are equally important. But if I had one wish, it'd be that the effects were larger. Three percentage points is statistically robust but thin. But the real contribution here is the way you tested things, not just the results.
Cite this project
@misc{gautam2026where,
title = {{Where Self-Knowledge Fails. Models Predict Their Own Choices Well, but Misreport the Ones That Concern Themselves}},
author = {Arpit Singh Gautam},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/where-selfknowledge-fails-models-predict-their-own-choices-well-but-misreport-the-ones-that-concern-themselves-u359}},
url = {https://apartresearch.com/sprints/projects/where-selfknowledge-fails-models-predict-their-own-choices-well-but-misreport-the-ones-that-concern-themselves-u359}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …