Manipulation Playground
Giles Edkins, Zoravur Singh, Christopher Berry · Team Playground
Submitted to AI Manipulation Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
As Large Language Models (LLMs) operate in multi-agent settings, understanding emergent manipulation and deception becomes increasingly important. While prior work focuses on LLMs manipulating humans, LLM-to-LLM dynamics remain understudied. We extend the Cheap Talk setup (Pham 2025) to create controlled scenarios where agents communicate under partially aligned or conflicting incentives. This framework enables observation of influence attempts, misrepresentation, information withholding, and covert signaling. We contextualize our approach within recent work on manipulation and deception, including APE and DeceptionBench, and offer preliminary observations about when models engage in manipulation. Early findings suggest that some agents alter their behavior depending on the capability of the victim, agents double-down on manipulative behavior in iterated games, and that manipulative tendencies arise without explicit prompting.
Reviews
Manipulation Playground effectively explores emergent strategic behaviors between LLMs using a multi-agent game setup. The graphs comparing different model pairs provide clear visual evidence of behaviors such as double-down manipulation and capability-dependent strategy shifts.
To improve, the project could include quantitative metrics for manipulative actions, more explicit definitions of what counts as manipulation, and additional experiments across a broader set of models or incentive structures. Overall, it is a well-executed exploratory study with promising insights into LLM-to-LLM interactions.
Cite this project
@misc{edkins2026manipulation,
title = {{Manipulation Playground}},
author = {Giles Edkins and Zoravur Singh and Christopher Berry},
year = {2026},
month = jan,
note = {Submitted to AI Manipulation Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/manipulation-playground-siwg}},
url = {https://apartresearch.com/sprints/projects/manipulation-playground-siwg}
}More from AI Manipulation Hackathon
- 1st placeView project: Who Does Your AI Serve? Manipulation By and Of AI Assistants
Who Does Your AI Serve? Manipulation By and Of AI Assistants
Cart Abandonment Issues 🛒
AI assistants can be both instruments and targets of manipulation. In our project, we investigated both directions across three studies. AI as Instrument: Operators can instruct AI to prioritise their interests at the …
- 2nd placeView project: Eliciting Deception on Generative Search Engines
Eliciting Deception on Generative Search Engines
Ardy
Large language models (LLMs) with web browsing capabilities are vulnerable to adversarial content injection—where malicious actors embed deceptive claims in web pages to manipulate model outputs. We investigate whether …
- 3rd placeView project: Cross-Linguistic Sycophancy in Frontier LLMs: A Benchmark Study
Cross-Linguistic Sycophancy in Frontier LLMs: A Benchmark Study
Talex
We developed a cross-linguistic sycophancy benchmark testing whether frontier AI models exhibit different manipulation behaviours across English, Japanese, and Bengali. Our results show significant language-dependent …