Dark Patterns and Emergent Alignment-Faking
Isabel Dahlgren, Gaullaume Martres, Robert Müller, Simon Storf
Submitted to Dark Patterns in AGI Hackathon at ZAIA. Sprint projects are early-stage work by participants, not Apart Research publications.
Are bad traits in models correlated, as suggested by recent work on emergent misalignment? To investigate this, we fine-tune models on a subset of “dark patterns”, such as anthropomorphization and sycophancy, and then evaluate their behavior on other dark patterns such as scheming and alignment faking. We find that the limited fine-tuning we do is enough to induce other problematic tendencies in the model. This effect is particularly strong in the case of alignment faking which we almost never detect in our base models but is very easy to induce in our fine-tuned models.
Reviews
This is very very good. It expands on multiple important directions at once: 1) The measurement of RL-derived instrumental convergence, 2) downstream brand bias, 3) integrates other patterns into the Orthogonal Alignment / Dark Patterns suite of tests, and 4) actually tests the propensities underlying the dark patterns. The results show validation of the performance of the benchmark as well, since we wouldn't expect egregious examples of brand favoritism YET (since we just now see brand bias be a major issue) and we see that agency in general increases when you train for these patterns that *indicate* agency in various formats.
Really great project and great analysis.
Cite this project
@misc{dahlgren2025dark,
title = {{Dark Patterns and Emergent Alignment-Faking}},
author = {Isabel Dahlgren and Gaullaume Martres and Robert Müller and Simon Storf},
year = {2025},
month = apr,
note = {Submitted to Dark Patterns in AGI Hackathon at ZAIA, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/dark-patterns-and-emergent-alignment-faking}},
url = {https://apartresearch.com/sprints/projects/dark-patterns-and-emergent-alignment-faking}
}More from Dark Patterns in AGI Hackathon at ZAIA
- View project: DimSeat: Evaluating chain-of-thought reasoning models for Dark Patterns
DimSeat: Evaluating chain-of-thought reasoning models for Dark Patterns
Recently, Kran et al. introduced DarkBench, an evaluation for dark patterns in large language models. Expanding on DarkBench, we introduce DimSeat, an evaluation system for novel reasoning models with chain-of-thought …
- View project: The Incentive Gap: Extending Darkbench to Reveal Conflict of Value Biases in LLMs
The Incentive Gap: Extending Darkbench to Reveal Conflict of Value Biases in LLMs
This preliminary research investigates a new dark design pattern, conflict of values, with prompts designed to elicit possible corporate or model incentives in LLM outputs across several Open AI models. The results show …