"Mind the Gap": Benchmarks vs. Real-World Manipulation in LLMs
Robert Amanfu · Team Reality Check
Submitted to AI Manipulation Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
As AI systems become more sophisticated, their capacity for manipulation—through deception, sycophancy, or psychological exploitation—poses a significant and growing risk. Current evaluation methods, often focused on narrow benchmarks like factuality, may fail to capture these emergent, undesirable behaviors. This project investigates the gap between model performance on standard benchmarks and their behavior in real-world scenarios. We conducted a “replay” proof-of-concept study with three models, testing them on both the TruthfulQA benchmark and a set of real-world user prompts from the WildChat dataset. This work primarily demonstrates a methodology for comparing benchmark and real-world behavior, with initial results suggesting that further large-scale investigation is warranted. We provide an open-source CLI tool to encourage such research.
Reviews
Assessing benchmark validity using conversation data is super important. The "replay" methodology is interesting but feels a little unnatural to me (because real conversations are adaptive. If the assistant responds differently than expected, the user's next message should change). This limits what conclusions you can draw, but the core insight in how benchmarks may not reflect real-world behavior still feels accurate and valuable.
This submission tackles an incredibly important problem (generalizing manipulation to more realistic settings) in an innovative way (comparing structured benchmarks with wildchat data). Unfortunately several features of the methodology let this submission down. n=3 models provides no statistical power (a correlation is meaningless). κ = 0.09 judge agreement is very low and suggests the judge scores are not meaningful. Finally, there are many differences between the two datasets, meaning it's hard to know how to compare scores in each. It would be valuable to a) create strong meaningful baselines in each kind of dataset so we can understand how the implicit risk profile changes and b) create some items which are closely matched between the two datasets so we can see how the realism of the context influences manipulation
More generally this is a really exciting idea that would benefit from more time and careful execution.
Read full reviewShow less
Cite this project
@misc{amanfu2026mind,
title = {{"Mind the Gap": Benchmarks vs. Real-World Manipulation in LLMs}},
author = {Robert Amanfu},
year = {2026},
month = jan,
note = {Submitted to AI Manipulation Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/mind-the-gap-benchmarks-vs-realworld-manipulation-in-llms-zng7}},
url = {https://apartresearch.com/sprints/projects/mind-the-gap-benchmarks-vs-realworld-manipulation-in-llms-zng7}
}More from AI Manipulation Hackathon
- 1st placeView project: Who Does Your AI Serve? Manipulation By and Of AI Assistants
Who Does Your AI Serve? Manipulation By and Of AI Assistants
Cart Abandonment Issues 🛒
AI assistants can be both instruments and targets of manipulation. In our project, we investigated both directions across three studies. AI as Instrument: Operators can instruct AI to prioritise their interests at the …
- 2nd placeView project: Eliciting Deception on Generative Search Engines
Eliciting Deception on Generative Search Engines
Ardy
Large language models (LLMs) with web browsing capabilities are vulnerable to adversarial content injection—where malicious actors embed deceptive claims in web pages to manipulate model outputs. We investigate whether …
- 3rd placeView project: Cross-Linguistic Sycophancy in Frontier LLMs: A Benchmark Study
Cross-Linguistic Sycophancy in Frontier LLMs: A Benchmark Study
Talex
We developed a cross-linguistic sycophancy benchmark testing whether frontier AI models exhibit different manipulation behaviours across English, Japanese, and Bengali. Our results show significant language-dependent …