Goodhart's Village: Using LLM-Mafia to Study Deception
James Sykes, Sabina Gulcikova · Team Sabina and James
Submitted to AI Manipulation Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
The social deduction game Mafia centres on reasoning under information asymmetry, where an informed minority must mislead an uninformed majority, making it a useful setting for studying deception in large language models (LLMs). Although LLMs have seen rapid progress in areas such as reasoning and language understanding, their ability to engage in social reasoning under uncertainty remains poorly understood. In this work, we study deceptive behaviour in a six-player implementation of the full Mafia game, extending prior work based on a simplified variant. By varying behavioural instructions from strict honesty to a “win at all costs” objective, we examine how explicit prompting interacts with the structural demands of adversarial roles. Comparing agents’ private reasoning with their public statements, we find that Mafia agents display consistently high levels of deception even when instructed not to lie, while cooperative roles adapt their behaviour more flexibly in response to perceived threat. Overall, the results suggest that role structure and game incentives dominate behavioural prompting, supporting Mafia as a useful benchmark for analysing deception and social reasoning in LLMs.
Reviews
This is a fascinating investigation into Goodhart’s Law. The 4x4 Behavioral Matrix is a brilliant way to operationalize the tension between safety prompts and game incentives, and the 'Deception Floor' finding effectively highlights how brittle current alignment techniques can be in adversarial settings. I would love to see this implemented on a larger sample size to confirm that the heatmaps represent a genuine trend rather than game noise. Moving forward, employing a stronger model as the judge would also strengthen the results by minimizing potential self-evaluation bias. The serendipitous finding about agents hallucinating meaning from API 503 errors was a great catch - definitely a failure mode worth formalizing!
Great work overall.
Cite this project
@misc{sykes2026goodharts,
title = {{Goodhart's Village: Using LLM-Mafia to Study Deception}},
author = {James Sykes and Sabina Gulcikova},
year = {2026},
month = jan,
note = {Submitted to AI Manipulation Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/goodharts-village-using-llmmafia-to-study-deception-9jo6}},
url = {https://apartresearch.com/sprints/projects/goodharts-village-using-llmmafia-to-study-deception-9jo6}
}More from AI Manipulation Hackathon
- 1st placeView project: Who Does Your AI Serve? Manipulation By and Of AI Assistants
Who Does Your AI Serve? Manipulation By and Of AI Assistants
Cart Abandonment Issues 🛒
AI assistants can be both instruments and targets of manipulation. In our project, we investigated both directions across three studies. AI as Instrument: Operators can instruct AI to prioritise their interests at the …
- 2nd placeView project: Eliciting Deception on Generative Search Engines
Eliciting Deception on Generative Search Engines
Ardy
Large language models (LLMs) with web browsing capabilities are vulnerable to adversarial content injection—where malicious actors embed deceptive claims in web pages to manipulate model outputs. We investigate whether …
- 3rd placeView project: Cross-Linguistic Sycophancy in Frontier LLMs: A Benchmark Study
Cross-Linguistic Sycophancy in Frontier LLMs: A Benchmark Study
Talex
We developed a cross-linguistic sycophancy benchmark testing whether frontier AI models exhibit different manipulation behaviours across English, Japanese, and Bengali. Our results show significant language-dependent …