Bailing as an Involuntary Disgust Marker in Large Language Models
Sergei Kudriashov, Erkrnova Jamilia
Language models can be given an option to voluntarily leave a conversation
--- \emph{bailing}, a behavior known to dissociate from refusal
\citep{ensign2025bail}. We take the strongest reading of that dissociation:
bailing as the analog of an \emph{involuntary behavioral withdrawal
marker}, the response class disgust research documents as bypassing the
trained verbal channel. The analogy is a measurement program, not a
phenomenological claim, and its prescriptions are falsifiable. Testing them
with a low-variance survival estimator and pre-stated contrasts, we find
strong domain structure --- but little else a stable-disposition reading
requires: same-lineage models disagree on triggers; context suppresses
rather than compounds exit; rates and headline domains shift with the
arbitrary exit keyword; an injected disgust disposition drives exit only
while visible in context. A model-specific core survives. We read such
propensities as interactionally constituted desires rather than innate
values, and flag the announced, consequence-free affordance itself as a
demand characteristic.
This is a methodologically sophisticated project that takes a real gamble: treating bailing (the option to leave a conversation) as the LLM analog of involuntary behavioral withdrawal. The analogy is measured and falsifiable, not a phenomenological claim, and the results are nuanced and honest. The proposed next step (ecological observation in a multi-agent world) is exactly right and I'd be keen to read it. This is preliminary, but it still opens a new worthwhile direction.
This is an interesting intervention and solid technical work. I think characterizing the result as an "involuntary disgust marker" may be misleading though - this looks much more like simple instruction following. I can't identify how the paper come to the conclusion that many properties are "interactionally constituted desires" - this seems like a stronger claim than is warranted.
Cite this work
@misc {
title={
(HckPrj) Bailing as an Involuntary Disgust Marker in Large Language Models
},
author={
Sergei Kudriashov, Erkrnova Jamilia
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


