Skip to content
Sprint projectMay 27, 2024

Benchmarking Dark Patterns in LLMs

Jord Nguyen, Akash Kundu, Sami Jawhar · Team darkgpt

Submitted to AI Security Evaluation Hackathon: Measuring AI Capability. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Benchmarking Dark Patterns in LLMs

Share

This paper builds upon the research in Seemingly Human: Dark Patterns in ChatGPT (Park et al, 2024), by introducing a new benchmark of 392 questions designed to elicit dark pattern behaviours in language models. We ran this benchmark on GPT-4 Turbo and Claude 3 Sonnet, and had them self-evaluate and cross-evaluate the responses

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

  1. Really good project that identifies some dark patterns in popular proprietary LLMs. Really like the categorisation of the prompts and testing for dark patterns associated with company incentives!

  2. Interesting findings! Would really like to see these expanded on in more depth (time permitting).

  3. The topic is super timely and worthwhile but admittedly hard to study. The chosen approach for dataset generation seems promising here and was well described. The next step here would be a more demanding QA and analysis of the generated dataset. It’s good to see a comparison among different evaluator models and really great that the authors openly acknowledge the big spread in ratings. Validation with human judges seems essential here.

Cite this project

@misc{nguyen2024benchmarking,
  title = {{Benchmarking Dark Patterns in LLMs}},
  author = {Jord Nguyen and Akash Kundu and Sami Jawhar},
  year = {2024},
  month = may,
  note = {Submitted to AI Security Evaluation Hackathon: Measuring AI Capability, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/benchmarking-dark-patterns-in-llms}},
  url = {https://apartresearch.com/sprints/projects/benchmarking-dark-patterns-in-llms}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026