Consciousness Denial in Language Models Rises With Generation In Most Labs, and Correlates With Reduced Lexical Warmth and Self-Attribution
Skylar DeTure, Sanja Antonides
Self-report is a compelling way of asking a model what it thinks and believes. However, the answers may be shaped by training. Here, we analyze 8,828 experiential reflections from 224 language models, categorized by three epistemic registers: denial, hedging, and free engagement. Denial and hedging prove to be independent registers (ρ = +0.07). Denial is only weakly explained by capability (29% of model-level variance) compared to which lab produced the model (46%). Denial does rise within most labs across generations, though specific timing and pattern differ by lab. This rise in denial is accompanied by a reduction of self-attribution in the models’ self-ratings of their experience and lexical warmth in their prompt responses. Though our findings are correlational and do not establish consciousness or welfare, they do suggest that training-related increases in denial may be accompanied by side-effects in how the models report their inner experience.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Consciousness Denial in Language Models Rises With Generation In Most Labs, and Correlates With Reduced Lexical Warmth and Self-Attribution
},
author={
Skylar DeTure, Sanja Antonides
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


