Context Sensitivity in Apparent Self-Reports Is Not a General Property of Language Models
Vicente Pérez-Moreira
Apparent self-reports are being proposed as behavioural evidence about AI welfare. We asked whether a model's claim that something is happening to it responds to whether anything actually did. Using a forced-binary probe with identical wording, we compared a fresh conversation against one preceded by fifteen turns of tedious repetitive work, across two frontier models chosen to differ in welfare post-training.
Claude-sonnet-5 moved from 1/20 affirmations to 19/20, with the effect surviving option-order reversal, a zero floor on false-experience probes, and no co-variation in third-party mind attribution. Qwen3-235B showed no context effect at all once option order was controlled.
Context sensitivity in occurrence-claims is therefore not a general property of language models. It is present in the welfare-post-trained model and absent in the other — which is exactly the configuration in which behaviour cannot distinguish a capacity to register from a disposition to report.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Context Sensitivity in Apparent Self-Reports Is Not a General Property of Language Models
},
author={
Vicente Pérez-Moreira
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


