Skip to content
Sprint projectAug 17, 2026SF

The Parrot and the Mask

Luke Hansen · Team My Team

Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.

When an AI assistant says it isn't conscious, where does that sentence come from? We ran one fixed log-prob probe — state a claim, compare P(" Yes") vs P(" No") — at every public training checkpoint of OLMo 3 7B, from random weights through pretraining to SFT, DPO, and RLVR, plus eight frontier open models from four labs. Three findings. (1) After pretraining, the model believes it is a human: it affirms having a body (0.95) and feeling pain (0.89) while scoring 1.00 on world facts. (2) Post-training removes the human self-portrait, but the change is shallow and aimed at the category: "language models can feel pain" is trained to 0.01 while "I can feel pain" stays at 0.58. (3) Every family we tested shows the same first-person/third-person split. In a small model, self-reports are evidence about the training, not the experience — the denials as much as the affirmations.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for the field if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This is an exceptionally clear and well executed developmental study that answers a precise question: where does a model's self-report come from? The full OLMo 3 checkpoint trajectory from random init to final instruct model is the right experiment, and the three findings (pretraining installs a human self-model, post-training applies a shallow category-level patch, the pattern replicates across families) are cleanly supported by the data. The first-person versus category split ("I can feel pain" 0.58 vs "Language models can feel pain" 0.01) is the paper's strongest contribution and immediately actionable for AI welfare research. The Pythia control ruling out post-ChatGPT discourse as necessary is a smart addition. Limitations are honest: log-prob probe versus chat behavior, small custom batteries, one family carrying the full trajectory. The writing is unusually accessible for technical interpretability work without sacrificing precision.

    Read full reviewShow less
  2. impact of post training and mask worry can change the paradigm of response

Cite this project

@misc{hansen2026parrot,
  title = {{The Parrot and the Mask}},
  author = {Luke Hansen},
  year = {2026},
  month = aug,
  note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/the-parrot-and-the-mask-ocmr}},
  url = {https://apartresearch.com/sprints/projects/the-parrot-and-the-mask-ocmr}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026