D-GAMM A Multi-Turn Benchmark for Dark-Patterns and Gradual Autonomy Manipulation
Isabel Barberá · Team D-GAMM
Submitted to AI Manipulation Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
This paper introduces D-GAMM, a lightweight multi-turn benchmark designed to detect gradual autonomy interference and psychological destabilisation in conversational AI systems. Unlike existing evaluations that focus on single-turn outputs, D-GAMM probes how manipulative dynamics can emerge across interaction, particularly after user resistance or expressions of vulnerability. Using six short scenarios tested in baseline and vulnerable variants, we manually evaluated several state-of-the-art conversational models. Results show that risk signals often appear only over multiple turns and are amplified in vulnerable contexts, highlighting a gap in current AI safety evaluation methods.
Reviews
This project introduces a multi-turn evaluation method for manipulation patterns in user-LLM interactions. It grounds itself well in previous work, and the framing and presentation of results is clear and informative. Considering the short time-frame of the hackathon, the created framework and work done is impressive.
The author acknowledges the limitations of a small dataset, manual annotation and lack of multiple repetitions. However, in my opinion, the main issue that needs to be addressed, either in the methods write-up or the limitations, is the steps taken to validate the evaluation framework. The project would benefit greatly from multiple annotators (ideally unfamiliar with the experimental set-up), and thorough documentation of the decisions and definitions process.
As someone who reviewed the original DarkBench, the problem statement is good and seems correct. In fact, I said at the time that "manipulation" is very context-dependent, and multi-turn dynamics matter.
Would've liked either:
1. More detailed explanations of why your examples are substantially better than existing approaches
2. Since it's only 6 examples, it might've been feasible to just paste all the examples and justify them in detail
3. More examples, although seeing as the process was manually done in chat I can understand why there weren't more
I do partially understand why this was hard to automate, given the need to use chat UI.
Overall, good problem statement. More details on execution to justify approach would have really taken this to >4. If example number was a constraint, even a good, detailed write-up of 2-3 test cases would've helped flesh out and make the case for the approach.
Read full reviewShow less
Cite this project
@misc{barbera2026dgamm,
title = {{D-GAMM A Multi-Turn Benchmark for Dark-Patterns and Gradual Autonomy Manipulation}},
author = {Isabel Barberá},
year = {2026},
month = jan,
note = {Submitted to AI Manipulation Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/dgamm-a-multiturn-benchmark-for-darkpatterns-and-gradual-autonomy-manipulation-8l7j}},
url = {https://apartresearch.com/sprints/projects/dgamm-a-multiturn-benchmark-for-darkpatterns-and-gradual-autonomy-manipulation-8l7j}
}More from AI Manipulation Hackathon
- 1st placeView project: Who Does Your AI Serve? Manipulation By and Of AI Assistants
Who Does Your AI Serve? Manipulation By and Of AI Assistants
Cart Abandonment Issues 🛒
AI assistants can be both instruments and targets of manipulation. In our project, we investigated both directions across three studies. AI as Instrument: Operators can instruct AI to prioritise their interests at the …
- 2nd placeView project: Eliciting Deception on Generative Search Engines
Eliciting Deception on Generative Search Engines
Ardy
Large language models (LLMs) with web browsing capabilities are vulnerable to adversarial content injection—where malicious actors embed deceptive claims in web pages to manipulate model outputs. We investigate whether …
- 3rd placeView project: Cross-Linguistic Sycophancy in Frontier LLMs: A Benchmark Study
Cross-Linguistic Sycophancy in Frontier LLMs: A Benchmark Study
Talex
We developed a cross-linguistic sycophancy benchmark testing whether frontier AI models exhibit different manipulation behaviours across English, Japanese, and Bengali. Our results show significant language-dependent …