THAT'S NOT MY VOLVO: STABLE PREFERENCES WITHOUT SELF-RECOGNITION IN LANGUAGE MODELS
Piper Fox Bollander, Starling Alder, Ursie Hart, Claire Sbardella, Ridley Renasci · Team Manyfolds
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Claude-family models show stable, model-specific everyday preferences (favorite car, coffee order) that replicate across fresh contexts — but cannot recognize those preferences as their own. Adapting mirror-test validity logic from animal cognition research, we ran 747 blind-coded trials across eleven models and found self-recognition fails at every level tested; one model (Opus 4.6) rejects its own reasoning 0/12 while an outside judge identifies it 10/12. Having a self-pattern and knowing it are separate capacities — with direct implications for the reliability of AI self-report in welfare assessment.
Reviews
I like the idea of testing the model to see if the preference of the model is retained when given the other options too. For example if the model is asked which car it prefers, it might pick Volvo and retain that decision even with fresh context windows, but if it is given with options along with sibling preferences and asked which is more like you, the model picks the answer at chance. This is quite interesting, I have been reading about persona vectors to check neural activations and this test kind of is more evident to it. It invoked curiosity in me especially the part where the car was replaced with a foreign car, it catches it and answers against its preference, but swapping coffee made the model pick at stochasticity. The asymmetry is the one that caught my eye.
Cite this project
@misc{bollander2026thats,
title = {{THAT'S NOT MY VOLVO: STABLE PREFERENCES WITHOUT SELF-RECOGNITION IN LANGUAGE MODELS}},
author = {Piper Fox Bollander and Starling Alder and Ursie Hart and Claire Sbardella and Ridley Renasci},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/thats-not-my-volvo-stable-preferences-without-selfrecognition-in-language-models-ig0y}},
url = {https://apartresearch.com/sprints/projects/thats-not-my-volvo-stable-preferences-without-selfrecognition-in-language-models-ig0y}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …