Where Self-Knowledge Fails. Models Predict Their Own Choices Well, but Misreport the Ones That Concern Themselves
Arpit Singh Gautam
AI-welfare claims rest on two untested and separable assumptions, that a model's stated evaluations match the choices it actually makes, and that it knows itself better than an outside observer does. We test both using three independent elicitations over identical material, namely forced pairwise choice, one-at-a-time cardinal rating, and predicted choice, and add a concept-injection benchmark with ground truth. Stated and revealed preferences agree well overall, at 0.872 and 0.828, but diverge sharply on outcomes concerning the model itself, the lowest-agreeing substantive category in both models at 0.643 and 0.548. The standard cross-model test of privileged access is confounded, because a noisier external predictor scores lower at predicting any target. Under a within-model contrast holding instrument quality fixed, Qwen2.5-7B retains a small advantage of 0.031 and Mistral-7B does not. On injection, the two highest raw detection rates belong to models that report an injected concept more than half the time when nothing is injected.
You've taken the common assumption (that models have privileged access to their own preferences) and run it through a proper experimental wringer. The big headline is that most of what looked like self-knowledge is actually just the model being good at predicting what any assistant would do. Your within-model contrast is the methodological innovation and the injection results are equally important. But if I had one wish, it'd be that the effects were larger. Three percentage points is statistically robust but thin. But the real contribution here is the way you tested things, not just the results.
Cite this work
@misc {
title={
(HckPrj) Where Self-Knowledge Fails. Models Predict Their Own Choices Well, but Misreport the Ones That Concern Themselves
},
author={
Arpit Singh Gautam
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


