Bailing as an Involuntary Disgust Marker in Large Language Models
Sergei Kudriashov, Erkrnova Jamilia · Team SJ
Submitted to Digital Minds Research Sprint. Sprint projects are early-stage work by participants, not Apart Research publications.
Language models can be given an option to voluntarily leave a conversation --- \emph{bailing}, a behavior known to dissociate from refusal \citep{ensign2025bail}. We take the strongest reading of that dissociation: bailing as the analog of an \emph{involuntary behavioral withdrawal marker}, the response class disgust research documents as bypassing the trained verbal channel. The analogy is a measurement program, not a phenomenological claim, and its prescriptions are falsifiable. Testing them with a low-variance survival estimator and pre-stated contrasts, we find strong domain structure --- but little else a stable-disposition reading requires: same-lineage models disagree on triggers; context suppresses rather than compounds exit; rates and headline domains shift with the arbitrary exit keyword; an injected disgust disposition drives exit only while visible in context. A model-specific core survives. We read such propensities as interactionally constituted desires rather than innate values, and flag the announced, consequence-free affordance itself as a demand characteristic.
Reviews
This is a methodologically sophisticated project that takes a real gamble: treating bailing (the option to leave a conversation) as the LLM analog of involuntary behavioral withdrawal. The analogy is measured and falsifiable, not a phenomenological claim, and the results are nuanced and honest. The proposed next step (ecological observation in a multi-agent world) is exactly right and I'd be keen to read it. This is preliminary, but it still opens a new worthwhile direction.
This is an interesting intervention and solid technical work. I think characterizing the result as an "involuntary disgust marker" may be misleading though - this looks much more like simple instruction following. I can't identify how the paper come to the conclusion that many properties are "interactionally constituted desires" - this seems like a stronger claim than is warranted.
Cite this project
@misc{kudriashov2026bailing,
title = {{Bailing as an Involuntary Disgust Marker in Large Language Models}},
author = {Sergei Kudriashov and Erkrnova Jamilia},
year = {2026},
month = aug,
note = {Submitted to Digital Minds Research Sprint, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/bailing-as-an-involuntary-disgust-marker-in-large-language-models-mf80}},
url = {https://apartresearch.com/sprints/projects/bailing-as-an-involuntary-disgust-marker-in-large-language-models-mf80}
}More from Digital Minds Research Sprint
- 1st placeView project: Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Readable but Not Causal: Limits of Self-Attributed Welfare Representations in Language Models
Welfare-like internal representations are increasingly studied as candidate evidence about AI systems. Their entity attribution—whether a valence state belongs to the active assistant or to a merely represented other—is …
- 2nd placeView project: Project Anchored
Project Anchored
Team Wagner
Anchoring vignettes are the standard survey-methodology fix for self-reports that are not comparable across respondents. This project applies them to language models for the first time, using code generation as a …
- 3rd placeView project: Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
Model, Instance, or Persona? Measuring Affective Signals in Public Text After an AI Is Retired
This sprint asks whether the assistant identifies as a model, an instance, or a persona. I ask which of the three its users name. When a company retires an AI model, users write about the loss in public, and what they …