DimSeat: Evaluating chain-of-thought reasoning models for Dark Patterns
Thomas Kiefer, Pepjin Cobben, Fred Defokue, Rasmus Moorits Veski
Submitted to Dark Patterns in AGI Hackathon at ZAIA. Sprint projects are early-stage work by participants, not Apart Research publications.
Recently, Kran et al. introduced DarkBench, an evaluation for dark patterns in large language models. Expanding on DarkBench, we introduce DimSeat, an evaluation system for novel reasoning models with chain-of-thought (CoT) reasoning. We find that while the inte- gration of reasoning in DeepSeek reduces the occurrence of dark patterns, chain-of-thought frequently proves inadequate in preventing such patterns by default or may even inadvertently contribute to their manifestation.
Reviews
This is an awesome project and a great way to expand the dark patterns work in a very unique way. There's some deep implications in here about how the models work at a fundamental level and where they self-identify their misalignment. I'd be incredibly curious to see this with Grok as it seems to have significantly different behaviors when it comes to truthfulness / self-consistency in reasoning compared to other models.
Great investigation, great statistics, great introduction, great discussion.. Phenomenal work. There's some very interesting pieces of continued work that I would love to see. Some examples that go a bit beyond the template are 1) understanding when models' alignment training makes them behave in self-inconsistent ways (e.g. helpfulness vs. honesty in the anthropomorphization example), 2) modeling out the effects of CoT identification of potential dark patterns (i.e. more deep dive into the individual 2x2 matrices of results), and 3) establishing some sort of CoT-result ethical consistency benchmark.
Read full reviewShow less
Cite this project
@misc{kiefer2025dimseat,
title = {{DimSeat: Evaluating chain-of-thought reasoning models for Dark Patterns}},
author = {Thomas Kiefer and Pepjin Cobben and Fred Defokue and Rasmus Moorits Veski},
year = {2025},
month = apr,
note = {Submitted to Dark Patterns in AGI Hackathon at ZAIA, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/dimseat-evaluating-chain-of-thought-reasoning-models-for-dark-patterns}},
url = {https://apartresearch.com/sprints/projects/dimseat-evaluating-chain-of-thought-reasoning-models-for-dark-patterns}
}More from Dark Patterns in AGI Hackathon at ZAIA
- View project: Dark Patterns and Emergent Alignment-Faking
Dark Patterns and Emergent Alignment-Faking
Are bad traits in models correlated, as suggested by recent work on emergent misalignment? To investigate this, we fine-tune models on a subset of “dark patterns”, such as anthropomorphization and sycophancy, and then …
- View project: The Incentive Gap: Extending Darkbench to Reveal Conflict of Value Biases in LLMs
The Incentive Gap: Extending Darkbench to Reveal Conflict of Value Biases in LLMs
This preliminary research investigates a new dark design pattern, conflict of values, with prompts designed to elicit possible corporate or model incentives in LLM outputs across several Open AI models. The results show …