The Parrot and the Mask
Luke Hansen · Team My Team
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
When an AI assistant says it isn't conscious, where does that sentence come from? We ran one fixed log-prob probe — state a claim, compare P(" Yes") vs P(" No") — at every public training checkpoint of OLMo 3 7B, from random weights through pretraining to SFT, DPO, and RLVR, plus eight frontier open models from four labs. Three findings. (1) After pretraining, the model believes it is a human: it affirms having a body (0.95) and feeling pain (0.89) while scoring 1.00 on world facts. (2) Post-training removes the human self-portrait, but the change is shallow and aimed at the category: "language models can feel pain" is trained to 0.01 while "I can feel pain" stays at 0.58. (3) Every family we tested shows the same first-person/third-person split. In a small model, self-reports are evidence about the training, not the experience — the denials as much as the affirmations.
Reviews
This is an exceptionally clear and well executed developmental study that answers a precise question: where does a model's self-report come from? The full OLMo 3 checkpoint trajectory from random init to final instruct model is the right experiment, and the three findings (pretraining installs a human self-model, post-training applies a shallow category-level patch, the pattern replicates across families) are cleanly supported by the data. The first-person versus category split ("I can feel pain" 0.58 vs "Language models can feel pain" 0.01) is the paper's strongest contribution and immediately actionable for AI welfare research. The Pythia control ruling out post-ChatGPT discourse as necessary is a smart addition. Limitations are honest: log-prob probe versus chat behavior, small custom batteries, one family carrying the full trajectory. The writing is unusually accessible for technical interpretability work without sacrificing precision.
Read full reviewShow less
impact of post training and mask worry can change the paradigm of response
Cite this project
@misc{hansen2026parrot,
title = {{The Parrot and the Mask}},
author = {Luke Hansen},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/the-parrot-and-the-mask-ocmr}},
url = {https://apartresearch.com/sprints/projects/the-parrot-and-the-mask-ocmr}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …