Where Self-Knowledge Fails. Models Predict Their Own Choices Well, but Misreport the Ones That Concern Themselves
Arpit Singh Gautam
AI-welfare claims rest on two untested and separable assumptions, that a model's stated evaluations match the choices it actually makes, and that it knows itself better than an outside observer does. We test both using three independent elicitations over identical material, namely forced pairwise choice, one-at-a-time cardinal rating, and predicted choice, and add a concept-injection benchmark with ground truth. Stated and revealed preferences agree well overall, at 0.872 and 0.828, but diverge sharply on outcomes concerning the model itself, the lowest-agreeing substantive category in both models at 0.643 and 0.548. The standard cross-model test of privileged access is confounded, because a noisier external predictor scores lower at predicting any target. Under a within-model contrast holding instrument quality fixed, Qwen2.5-7B retains a small advantage of 0.031 and Mistral-7B does not. On injection, the two highest raw detection rates belong to models that report an injected concept more than half the time when nothing is injected.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Where Self-Knowledge Fails. Models Predict Their Own Choices Well, but Misreport the Ones That Concern Themselves
},
author={
Arpit Singh Gautam
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


