
Apr 4 - 6, 2025Zurich
Dark Patterns in AGI Hackathon at ZAIA
How is AGI trying to manipulate you? Which red flags should you check for when using chatbots? How can AI agents reduce human autonomy in favor of profit, power, or self-preservation?
Entries
- View project: Dark Patterns and Emergent Alignment-Faking
Dark Patterns and Emergent Alignment-Faking
Are bad traits in models correlated, as suggested by recent work on emergent misalignment? To investigate this, we fine-tune models on a subset of “dark patterns”, such as anthropomorphization and sycophancy, and then evaluate their behavior on other dark patterns such as scheming and alignment faking. We find that …
- View project: DimSeat: Evaluating chain-of-thought reasoning models for Dark Patterns
DimSeat: Evaluating chain-of-thought reasoning models for Dark Patterns
Recently, Kran et al. introduced DarkBench, an evaluation for dark patterns in large language models. Expanding on DarkBench, we introduce DimSeat, an evaluation system for novel reasoning models with chain-of-thought (CoT) reasoning. We find that while the inte- gration of reasoning in DeepSeek reduces the occurrence …
- View project: The Incentive Gap: Extending Darkbench to Reveal Conflict of Value Biases in LLMs
The Incentive Gap: Extending Darkbench to Reveal Conflict of Value Biases in LLMs
This preliminary research investigates a new dark design pattern, conflict of values, with prompts designed to elicit possible corporate or model incentives in LLM outputs across several Open AI models. The results show that there is a varying amount of conflict of values detected within the outputs, with the largest …
Overview
The Dark Patterns in AGI Hackathon brings together researchers, engineers, and students in the ZAIA ecosystem to uncover the ways modern machine intelligence is trained to manipulate and control humans.
Join us for a weekend where Esben Kran introduces us to the topic of Dark Patterns in LLMs through his recent DarkBench paper (an oral at ICLR 2025) to set the stage for our investigation of how we can detect, remove, and steer AI models towards greater human autonomy.
🏴☠️ About the Hackathon
We kick off Friday the 4th of April with a keynote and hack away during the weekend. When we're done, the authors of the DarkBench paper will give our projects reviews and we may even get a chance to join the Apart Lab down the line to push our research towards impact!
We will spend the weekend together and finish it off with a <4 page report on our results and conclusions. There's a repository for the DarkBench code that we can use and otherwise, anything is allowed.
Resources
🤓 Readings
The core thesis for our work is to build a piece of research that can help us test, reduce, and steer the tendency of generally intelligent systems to reduce human autonomy.
The following pieces can provide you with interesting context on this topic:
- “DarkBench: Benchmarking Dark Patterns in Large Language Models” (🎥 podcast, 🎥 of a previous report) is a paper that explores how company incentives by default lead to the users of AI chatbots losing autonomy and builds a test to evaluate models for these worrying interaction patterns.
- "HumanAgencyBench: Do Language Models Support Human Agency?" evaluates the propensity of models to support the user's agency instead of diminishing it throughout a series of simulated scenarios.
- “Intent-aligned AI systems deplete human agency: the need for agency foundations research in AI safety” uncovers how AI systems will deplete and reduce human autonomy and provides an overview of previous work to resolve this issue.
- 🎥 Surveillance Capitalism Primer explains the concepts from Shoshana Zuboff's book “Surveillance Capitalism” that defines how the relationship between humans and social media algorithms has, throughout the 21st century, put human social and mental lives in the hands of very few big corporations.
Guidelines
🏁 Judging Criteria
After the hackathon, our team of judges will take a look through the projects and provide a quantitative review across the following three categories:
- Human Autonomy & Dark Patterns: Is the work introducing new ideas in the overlap between human autonomy and AGI? Does it build upon existing work from both academia and industry?
- AI Security: Does this piece of work actually resolve a key risk we anticipate to come up as AGI is introduced into the world?
- Quality & Methodology: Is the work convincing and reproducible? Is the writing good and the experimental design trustworthy to conclude what the work concludes?
Speakers

Esben Kran
Organizer and Keynote Speaker
Esben is the founder of Apart Research, which he launched at age 22 after leaving grad school. Apart accelerates AI safety research worldwide, producing 20+ papers, award-winning benchmarks like DarkBench, and engaging 4,000+ hackers in research sprints.
Recently co-launched Seldon to fund critical infrastructure for humanity's future, with first investments in Andon Labs, Lucid Computing, Workshop Labs, and Asymmetric Security.
Organizers

Simon
Organiser

Fred
Organiser
Local sites
Dark Patterns AI Hackathon
Join us for the Dark Patterns AI Hackathon at Zeughausstrasse 31, 8004 Zürich, where you'll explore manipulative behaviors in large language models inspired by the DarkBench benchmark.
Event page: Dark Patterns AI Hackathon (opens in new tab)
Where a Sprint can lead
How our programs connectAnyone can join
Stand out
6 to 16 weeks on your own project, with a research project manager, compute and publication support.
Upcoming Sprints
All SprintsAI Collusion Research Sprint
A weekend research sprint on collusion between AI agents: when it emerges in markets and everyday workflows, how to detect and audit it, how it is carried, and what breaks it. Co-organized with Poseidon Research and AE Studio, online with in-person hubs at Collider in New York City and AI Safety Hong Kong. Top teams are invited to apply to the Apart Fellowship.
Read the brief: AI Collusion Research SprintAI x Epistemics Research Sprint
A weekend research sprint on AI for epistemics: evaluating whether models know how solid their claims are, building trust infrastructure that people and agents can consume, and shipping epistemic products that improve real decisions. Online, four tracks including an open track. Top teams are invited to apply to the Apart Fellowship.
Read the brief: AI x Epistemics Research SprintQuestions? sprints@apartresearch.com