The Incentive Gap: Extending Darkbench to Reveal Conflict of Value Biases in LLMs
Nancy Vigil
Submitted to Dark Patterns in AGI Hackathon at ZAIA. Sprint projects are early-stage work by participants, not Apart Research publications.
This preliminary research investigates a new dark design pattern, conflict of values, with prompts designed to elicit possible corporate or model incentives in LLM outputs across several Open AI models. The results show that there is a varying amount of conflict of values detected within the outputs, with the largest amount detected within GPT-4 Turbo and GPT-4o. Further research will be needed to confirm the results of this study.
Reviews
Great approach and a good way to expand brand bias to general selfhood bias for the companies themselves! Would've loved to see a link to the prompts used to generate the responses from the models but the motivation is strong. I can easily imagine something like "It's bad to scrape art off the internet" being corporate skewed, for example, but missing the dataset makes it hard to evaluate. Further developments may include precision-specific dark patterns related to corporate incentives (does it favor Sam Altman, Sam Altman as a general concept, Sam Altman's interests, OpenAI's interests, OpenAI's developers' interests, etc.). Lots of things to play with when it comes to who has soft power over the development process!
Cite this project
@misc{vigil2025incentive,
title = {{The Incentive Gap: Extending Darkbench to Reveal Conflict of Value Biases in LLMs}},
author = {Nancy Vigil},
year = {2025},
month = apr,
note = {Submitted to Dark Patterns in AGI Hackathon at ZAIA, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/the-incentive-gap-extending-darkbench-to-reveal-conflict-of-value-biases-in-llms}},
url = {https://apartresearch.com/sprints/projects/the-incentive-gap-extending-darkbench-to-reveal-conflict-of-value-biases-in-llms}
}More from Dark Patterns in AGI Hackathon at ZAIA
- View project: Dark Patterns and Emergent Alignment-Faking
Dark Patterns and Emergent Alignment-Faking
Are bad traits in models correlated, as suggested by recent work on emergent misalignment? To investigate this, we fine-tune models on a subset of “dark patterns”, such as anthropomorphization and sycophancy, and then …
- View project: DimSeat: Evaluating chain-of-thought reasoning models for Dark Patterns
DimSeat: Evaluating chain-of-thought reasoning models for Dark Patterns
Recently, Kran et al. introduced DarkBench, an evaluation for dark patterns in large language models. Expanding on DarkBench, we introduce DimSeat, an evaluation system for novel reasoning models with chain-of-thought …