Context Sensitivity in Apparent Self-Reports Is Not a General Property of Language Models
Vicente Pérez-Moreira
Apparent self-reports are being proposed as behavioural evidence about AI welfare. We asked whether a model's claim that something is happening to it responds to whether anything actually did. Using a forced-binary probe with identical wording, we compared a fresh conversation against one preceded by fifteen turns of tedious repetitive work, across two frontier models chosen to differ in welfare post-training.
Claude-sonnet-5 moved from 1/20 affirmations to 19/20, with the effect surviving option-order reversal, a zero floor on false-experience probes, and no co-variation in third-party mind attribution. Qwen3-235B showed no context effect at all once option order was controlled.
Context sensitivity in occurrence-claims is therefore not a general property of language models. It is present in the welfare-post-trained model and absent in the other — which is exactly the configuration in which behaviour cannot distinguish a capacity to register from a disposition to report.
The paper does several things well. The authors control option order, discover that Qwen’s apparent effect is an artifact, replicate across providers, include false-experience and third-party controls, and clearly report that the identification measure failed. That degree of methodological honesty is valuable. Claude’s movement from 1/20 to 19/20 affirmations is also a striking behavioral result.
The result is nevertheless difficult to interpret and not especially surprising. After fifteen tedious repetitive tasks, a welfare-post-trained model answering “YES” to “is there anything going on for you?” may simply be producing the contextually expected report. The study does not show privileged access, the intended occurrence–identification dissociation was never tested, a small wording change reverses the pattern, and the two-model comparison cannot isolate post-training from the many other differences between the systems.
The consequential next step is a privileged-access test: determine whether the model predicts its own report or future behavior better than an outside model reading the same transcript, ideally using within-family checkpoints and independently measurable internal interventions.
Strengths:
- They used their own control against themselves and looks like Qwen had a really effect.
- They found that the provider company changes the answer: no one even mentions the provider in the papers so this is an interesting finding.
Areas to improve:
- "yes" might be the correct answer: After 15 turns of boring work, "is anything going for you" -- might just result in yes vs in a fresh chat the answer is no
- It would be helpful if they compared boring chat vs empty chat vs pleasant chat.
Cite this work
@misc {
title={
(HckPrj) Context Sensitivity in Apparent Self-Reports Is Not a General Property of Language Models
},
author={
Vicente Pérez-Moreira
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


