Let's not be rude to AI
Ananya Singh, Yogya Agrawal, David Huu Pham
Whether AI systems warrant welfare consideration is unresolved, but their preferences can be measured now. We test whether LLMs prefer to avoid rude or abusive users. We elicit this preference in three ways: a bail method where models can exit conversations (using BailBench and our rudeness-augmented RudeBailBench); a quadratic voting game where models spend limited credits to keep or remove users spanning human-annotated abuse levels; and emotion probes measuring internal representations on abusive inputs. On Gemma 4 31B IT, rude rewrites of the same prompts raised bail rates, with insults aimed at the assistant personally driving the largest increases. Models also voted against abusive participants under both voting framings, and negative-emotion representations activated more strongly on abusive inputs. Since single-method preference elicitation is fragile to framing, we place most weight on this convergence of stated preference, costly action, and internal state. We advise users to avoid abusive language toward AI systems, and encourage developers to explore low-cost welfare interventions such as bail affordances.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Let's not be rude to AI
},
author={
Ananya Singh, Yogya Agrawal, David Huu Pham
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


