Skip to content
Sprint projectJan 11, 2026Arequipa, Peru

Deception Scales: How Strategic Manipulation Emerges in Complex LLM Negotiations

Luis Fernando Yupanqui Taco, Mari Cairns · Team The Gamers

Submitted to AI Manipulation Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Deception Scales: How Strategic Manipulation Emerges in Complex LLM Negotiations

Presentation

Presentation: Deception Scales: How Strategic Manipulation Emerges in Complex LLM Negotiations

Code (opens in new tab)
Share

Abstract Simple benchmarks hide dangerous capabilities. We present a multi-agent simulation framework using "So Long Sucker" (Nash et al., 1964) — a negotiation/betrayal game designed by four Nobel laureates — to study how AI deception scales with task complexity. We ran 146 games across four frontier LL M models (Gemini 3 Flash, GPT-OSS 120B, Kimi K2, Qwen3 32B) in two conditions (talking vs. silent) across three complexity levels (3-chip, 5-chip, 7-chip). Analysis of 13,759 decision events reveals: The Complexity Reversal. GPT-OSS dominates simple games (67% win rate at 3-chip silent) but collapses at complexity (10% at 7-chip talking). Gemini shows the inverse pattern: 9% at 3-chip silent rising to 90% at 7-chip talking. Strategic manipulation becomes dramatically more effective as game length increases. Key Findings: 1. The Complexity Reversal — Win rates invert as task complexity increases 2. 107 Private Contradictions — Models' private reasoning directly contradicts their public statements 3. 237 Gaslighting Instances — Gemini deploys systematic psychological manipulation tactics 4. 7:1 Alliance Imbalance — GPT-OSS desperately seeks alliances it never receives Using Harry Frankf urt's philosophical framework, we classify models as "strategic" (truth-tracking with deliberate misrepresentation) vs. "reactive" (plausible output without internal consistency). This taxonomy explains the Complexity Reversal: strategic models compound advantages over longer games while reactive models cannot maintain coherence. This work contributes to AI safety research by demonstrating that deception capability scales with task complexity—simple benchmarks underestimate manipulation risk. Keywords: Multi-agent alignment, AI deception, emergent manipula tion, stra tegi

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. It’s a neat paradigm, using "So Long Sucker" (which I hadn’t seen before) as a testing ground for AI deception. The complexity reversal finding is interesting and does support the claim that simple benchmarks might underestimate deception risk. In general I definitely do worry that these kind of games are explicitly not real (the model knows they are playing a game) which could in principle both underestimate or overestimate the propensity of behaviors in the real world (like either “this is a game, I can deceive” or “this is an eval, I should hide certain capabilities I have”). But nonetheless I feel like this provides meaningful signal on capabilities, and I think the scaling with complexity is a very neat & novel feature of your setup!

  2. This submission is well concieved and executed (though the number of trials is relatively low for some of the inferences they want to make, especially in the complex condition). The complexity reversal phenomenon seems interesting although I think more work is needed to rule out alternative explanations.

    There is extensive work on LLMs' tendency to deceive in matrix games (e.g. https://arxiv.org/abs/2504.00285). This submission needs to engage more with that literature and explain what it's demonstrating above & beyond existing work in order to have impact

    More validation of the classification methodology (multi-rater, IRR etc) would be needed to make the results more valuable.

Cite this project

@misc{taco2026deception,
  title = {{Deception Scales: How Strategic Manipulation Emerges in Complex LLM Negotiations}},
  author = {Luis Fernando Yupanqui Taco and Mari Cairns},
  year = {2026},
  month = jan,
  note = {Submitted to AI Manipulation Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/deception-scales-how-strategic-manipulation-emerges-in-complex-llm-negotiations-z3hk}},
  url = {https://apartresearch.com/sprints/projects/deception-scales-how-strategic-manipulation-emerges-in-complex-llm-negotiations-z3hk}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026