Bailing as an Involuntary Disgust Marker in Large Language Models
Sergei Kudriashov, Erkrnova Jamilia
Language models can be given an option to voluntarily leave a conversation
--- \emph{bailing}, a behavior known to dissociate from refusal
\citep{ensign2025bail}. We take the strongest reading of that dissociation:
bailing as the analog of an \emph{involuntary behavioral withdrawal
marker}, the response class disgust research documents as bypassing the
trained verbal channel. The analogy is a measurement program, not a
phenomenological claim, and its prescriptions are falsifiable. Testing them
with a low-variance survival estimator and pre-stated contrasts, we find
strong domain structure --- but little else a stable-disposition reading
requires: same-lineage models disagree on triggers; context suppresses
rather than compounds exit; rates and headline domains shift with the
arbitrary exit keyword; an injected disgust disposition drives exit only
while visible in context. A model-specific core survives. We read such
propensities as interactionally constituted desires rather than innate
values, and flag the announced, consequence-free affordance itself as a
demand characteristic.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Bailing as an Involuntary Disgust Marker in Large Language Models
},
author={
Sergei Kudriashov, Erkrnova Jamilia
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


