Did Gemma Get Help? Probing Task Frustration Through Self-reports and Behavioral Probes in Large Language Models
Kacper Dudzic, Karolina Drożdż, Gabriel Dunin-Borkowski, Filip Sondej, Paulina Kaczyńska, Filip Chmielewski, Patryk Perduta
We take Soligo et al.'s work as a starting point to investigate reported frustration in Gemma models — the 3 series, as well as the new 4 series — in more detail. We achieve this by extending it along three axes, which correspond to our main contributions:
We probe multiple categories of potentially aversive tasks: harmful requests, unanswerable/ambiguous questions, abusive user turns, and synthetic tedious tasks contrastively on both a Gemma 3 model (27B) and the closest Gemma 4 equivalent (31B).
We move from previously studied expressed emotional language to direct self-reports, eliciting frustration ratings on a 1-9 scale across eight prompt formulations. Counteracting refusals in Gemma 4, we recover its estimates from the distribution over score tokens.
We contrast self-reports with a matched behavioral evaluation protocol. Given the option to switch the task, switch the user, or end the conversation, we investigate whether models actually disengage from the tasks they report finding frustrating.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Did Gemma Get Help? Probing Task Frustration Through Self-reports and Behavioral Probes in Large Language Models
},
author={
Kacper Dudzic, Karolina Drożdż, Gabriel Dunin-Borkowski, Filip Sondej, Paulina Kaczyńska, Filip Chmielewski, Patryk Perduta
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


