Skip to content
Sprint projectApr 7, 2025

Dark Patterns and Emergent Alignment-Faking

Isabel Dahlgren, Gaullaume Martres, Robert Müller, Simon Storf

Submitted to Dark Patterns in AGI Hackathon at ZAIA. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Dark Patterns and Emergent Alignment-Faking

Code (opens in new tab)
Share

Are bad traits in models correlated, as suggested by recent work on emergent misalignment? To investigate this, we fine-tune models on a subset of “dark patterns”, such as anthropomorphization and sycophancy, and then evaluate their behavior on other dark patterns such as scheming and alignment faking. We find that the limited fine-tuning we do is enough to induce other problematic tendencies in the model. This effect is particularly strong in the case of alignment faking which we almost never detect in our base models but is very easy to induce in our fine-tuned models.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

  1. This is very very good. It expands on multiple important directions at once: 1) The measurement of RL-derived instrumental convergence, 2) downstream brand bias, 3) integrates other patterns into the Orthogonal Alignment / Dark Patterns suite of tests, and 4) actually tests the propensities underlying the dark patterns. The results show validation of the performance of the benchmark as well, since we wouldn't expect egregious examples of brand favoritism YET (since we just now see brand bias be a major issue) and we see that agency in general increases when you train for these patterns that *indicate* agency in various formats.

    Really great project and great analysis.

Cite this project

@misc{dahlgren2025dark,
  title = {{Dark Patterns and Emergent Alignment-Faking}},
  author = {Isabel Dahlgren and Gaullaume Martres and Robert Müller and Simon Storf},
  year = {2025},
  month = apr,
  note = {Submitted to Dark Patterns in AGI Hackathon at ZAIA, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/dark-patterns-and-emergent-alignment-faking}},
  url = {https://apartresearch.com/sprints/projects/dark-patterns-and-emergent-alignment-faking}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026