Distress Representations in Language Models Are Referent-Specific
Ayodeji Adesegun , Moyinoluwa Ogunjobi
AI welfare evaluations read internal “distress” directions as evidence about a model’s condition, but every published battery confounds it with the sentiment of the text and with distress attributed to others. Holding the event fixed, we vary only its referent: the model itself, another language model, or a fictional android, crossed with valence, so the referent cancels within each frame. Across six open-weight models (0.5B–14B), self- and other-referential distress vectors separate: Δ peaks at +0.348 (95% CI [+0.284, +0.403]), significant in 39 of 40 cells against reliability ceilings of 0.894–0.972. Self-reported valence tracks the self-specific residual (β = −0.33) as strongly as the shared component (β = −0.35). A second-person control with a non-self referent shows the separation is referential, not grammatical. Referential controls should be standard in welfare evaluations; none currently use them.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Distress Representations in Language Models Are Referent-Specific
},
author={
Ayodeji Adesegun , Moyinoluwa Ogunjobi
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


