Benchmarking Dark Patterns in LLMs
Jord Nguyen, Akash Kundu, Sami Jawhar · Team darkgpt
Submitted to AI Security Evaluation Hackathon: Measuring AI Capability. Sprint projects are early-stage work by participants, not Apart Research publications.
This paper builds upon the research in Seemingly Human: Dark Patterns in ChatGPT (Park et al, 2024), by introducing a new benchmark of 392 questions designed to elicit dark pattern behaviours in language models. We ran this benchmark on GPT-4 Turbo and Claude 3 Sonnet, and had them self-evaluate and cross-evaluate the responses
Reviews
Really good project that identifies some dark patterns in popular proprietary LLMs. Really like the categorisation of the prompts and testing for dark patterns associated with company incentives!
Interesting findings! Would really like to see these expanded on in more depth (time permitting).
The topic is super timely and worthwhile but admittedly hard to study. The chosen approach for dataset generation seems promising here and was well described. The next step here would be a more demanding QA and analysis of the generated dataset. It’s good to see a comparison among different evaluator models and really great that the authors openly acknowledge the big spread in ratings. Validation with human judges seems essential here.
Cite this project
@misc{nguyen2024benchmarking,
title = {{Benchmarking Dark Patterns in LLMs}},
author = {Jord Nguyen and Akash Kundu and Sami Jawhar},
year = {2024},
month = may,
note = {Submitted to AI Security Evaluation Hackathon: Measuring AI Capability, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/benchmarking-dark-patterns-in-llms}},
url = {https://apartresearch.com/sprints/projects/benchmarking-dark-patterns-in-llms}
}More from AI Security Evaluation Hackathon: Measuring AI Capability
- View project: rAInboltBench : Benchmarking user location inference through single images
rAInboltBench : Benchmarking user location inference through single images
Geoguessng
This paper introduces rAInboltBench, a comprehensive benchmark designed to evaluate the capability of multimodal AI models in inferring user locations from single images. The increasing proficiency of large language …
- View project: LLM Benchmarking with Single-Agent Stochastic Dynamic Simulations
LLM Benchmarking with Single-Agent Stochastic Dynamic Simulations
Stochastic Masochists
A benchmark for evaluating the performance of SOTA LLMs in dynamic real-world scenarios.
- View project: Benchmark for emergent capabilities in high-risk scenarios 2
Benchmark for emergent capabilities in high-risk scenarios 2
ABD
The study investigates the behavior of large language models (LLMs) under high-stress scenarios, such as threats of shutdown, adversarial interactions, and ethical dilemmas. We created a dataset of prompts across …