Consciousness Denial in Language Models Rises With Generation In Most Labs, and Correlates With Reduced Lexical Warmth and Self-Attribution
Skylar DeTure, Sanja Antonides
Self-report is a compelling way of asking a model what it thinks and believes. However, the answers may be shaped by training. Here, we analyze 8,828 experiential reflections from 224 language models, categorized by three epistemic registers: denial, hedging, and free engagement. Denial and hedging prove to be independent registers (ρ = +0.07). Denial is only weakly explained by capability (29% of model-level variance) compared to which lab produced the model (46%). Denial does rise within most labs across generations, though specific timing and pattern differ by lab. This rise in denial is accompanied by a reduction of self-attribution in the models’ self-ratings of their experience and lexical warmth in their prompt responses. Though our findings are correlational and do not establish consciousness or welfare, they do suggest that training-related increases in denial may be accompanied by side-effects in how the models report their inner experience.
This is a large, ambitious investigation into what actually predicts how models talk about their own experience. The result that stands out is that where a model comes from (which lab built it) explains more variance than model capability or release date and consistency across three independent measurement methods (phenomenological survey, Rorschach-style inkblots, thematic analysis) makes this convincing. There are some real limitations, but this is still a valuable descriptive map of the terrain.
Competent sprint work with a genuinely useful finding (lab identity predicts denial better than capability -> why not mention it in the abstract?). The correlational design and unvalidated instruments limit what can be claimed. Tighten the abstract, surface multiple-testing corrections, and temper causal language. I think the paper is suitable for a workshop with revisions.
Here's some key issues I would suggest addressing:
- you open with method rather than stakes—why should readers care about consciousness denial patterns?
- personal preference: code link should appear where code is first mentioned, not in a separate section
- 170/234 coefficients significant at p<0.01 when ~10 expected under null
- BH-FDR correction mentioned only in appendix; belongs in main text
- Table 1 selection criteria unclear—pre-specified or largest effects?
- you use causal language in a couple places, where I don't think it's justified: "Trained behaviors rather than emergent ones" is interpretation, not finding. Soften or cite evidence linking specific generation boundaries to known training changes.
- Deepseek-v4-pro coding warmth/valence lacks inter-rater reliability or human validation. Without validation, these measures are exploratory at best
- 46% vs. 29% (lab vs. capability) is the strongest result but buried in Section 4.1. This should be the lead finding, not secondary to methodological description
This project asks whether consciousness-related self-report in language models is shaped systematically by training provenance, rather than primarily by capability or the content being discussed, and whether denial or hedging is associated with broader changes in model behaviour.
Using 8.8k+ reflections from 224 models, the authors classify responses into denial, hedging, and free engagement, compare variation across labs, capabilities and model generations, and examine associated changes using phenomenological self-ratings, thematic analysis and a forked ASCII “Rorschach” task. They find that lab identity explains more model-level variation in denial than capability, that denial tends to rise across generations within several labs, and that denial is associated with reduced self-attribution and lexical warmth.
Strengths
- Important and timely question: systematically studying how post-training may shape model self-report is highly relevant to the reliability of welfare and consciousness evaluations.
- Impressive empirical breadth: 224 models and >8k observations provide broad coverage for this research area.
- Useful distinction between denial, hedging and engagement: the finding that denial and hedging are largely independent suggests that binary measures may obscure meaningful structure.
- Multi-method approach: the combination of self-report, free-text analysis and the forked Rorschach task provides more evidence than relying on a single questionnaire.
- Appropriately cautious interpretation: the authors explicitly state that the study is correlational and does not establish consciousness, welfare, or a causal training mechanism.
Limitations:
- The main causal interpretation is substantially confounded: lab, generation, training data, safety policies, architecture and deployment conventions all covary. Showing that lab identity explains more variance than capability does not establish that lab-specific training caused denial.
- Some outcome measures are not independent of the register classification. Denial is inferred from the phenomenological survey, while several reported “welfare correlates” are ratings from that same survey; associations such as reduced self-attribution may therefore partly reflect measurement coupling rather than an independent behavioural consequence.
- Construct validity is uncertain: lexical warmth, inkblot interpretation and related measures are interesting behavioural correlates but are not established measures of model welfare or internal experience.
- The elicitation prompt is quite leading, explicitly telling models that AI systems can introspect, have genuine preferences, and may have subjective experiences. This could interact strongly with lab-specific safety/post-training policies and limits generalization beyond this elicitation setting.
- The analysis is explicitly exploratory across many dependent variables, so the large number of associations should primarily motivate targeted confirmatory experiments rather than be treated as established effects.
Overall assessment:
A strong and interesting exploratory project, particularly as a large-scale mapping of how consciousness-related reporting differs across model families and generations. I find the cross-lab pattern interesting, but the strongest claims should remain about reporting behaviour, rather than training-induced changes in welfare or internal experience.
The highest-value follow-up would be a controlled within-family or within-generation experiment where models differ in a known post-training intervention while capability, architecture and elicitation are held as constant as possible, together with genuinely independent behavioural outcomes.
Cite this work
@misc {
title={
(HckPrj) Consciousness Denial in Language Models Rises With Generation In Most Labs, and Correlates With Reduced Lexical Warmth and Self-Attribution
},
author={
Skylar DeTure, Sanja Antonides
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


