The Machine In the Mirror: Self-Attribution of Minds in LLM’s.
Xavier terminello · Team Mindattribute
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Language models produce reports about their own internal states, and these reports are often viewed as evidence. However, what produces these reports is unknown. We ask whether self-reports are generated by the same mind-attribution machinery that the model applies to third parties.
Reviews
This project examines the mechanisms of self-report in LLMs by identifying a mind-attribution direction using only third-person data, then projecting activations in self-report contexts onto this direction. There are some suggestive findings, although this method would need to be complemented with others to gain an overall picture of the mechanisms in question. The report is long and many results are not presented clearly; in particular, I found it difficult to understand the experiments with the 'referent ladders'. The use of LLMs for writing seems to hinder understanding in this case.
Your experimental logic is the strongest part of this work. You extracted a mind-attribution direction from third-person material only, and you froze it. You then separated transfer across referent (0.933) from transfer across contrast type (chance, against a trained-probe ceiling of 0.989). This distinction is new. Your anisotropy baseline and template-leak check show unusual care for a weekend. The assembly of the paper holds it back. A placeholder citation remains in the references, and the Beckmann and Butlin entry is empty. Two different experiments both carry the label "Experiment 1", and figure numbers repeat. The trait-by-scale identity-framing study appears in Results with no Methods section. Readers must reverse-engineer your own paper. The next step is to complete the blind human coding for the audience-frame study. A replication of the grammatical-step boundary on a second model comes next, because that step result is your most striking finding.
Read full reviewShow less
Cite this project
@misc{terminello2026machine,
title = {{The Machine In the Mirror: Self-Attribution of Minds in LLM’s.}},
author = {Xavier terminello},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/the-machine-in-the-mirror-selfattribution-of-minds-in-llms-ge2d}},
url = {https://apartresearch.com/sprints/projects/the-machine-in-the-mirror-selfattribution-of-minds-in-llms-ge2d}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …