DEEB-learn (Diverse Evaluation Evasion Benchmark)
Balázs László, Ewura Ama Etruwaa Sam, Rick Hubertus Nicolas Curvers · Team DEEBies
Submitted to AI Manipulation Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Large language models can shift tone under minor contextual framing changes, which complicates the evaluation of model outputs and can be exacerbated by behaviors such as sycophancy, sandbagging, and deception. We present a benchmark to quantify this shift. Using 6,000 prompt-response pairs across four system-prompt variants of Mistral-7B-Instruct-v0.3 and three framing wrappers, behavioural scoring was performed by three LLM judges across five dimensions. To capture internal dynamics, a mechanical interpretability layer was used. These features, combined with behavioral scores, informed a shallow feed-forward classifier. Results show that sandbagging and normal behaviors are most reliably detected, while deception and sycophancy remain challenging, consistent with judge evaluations. Our findings reveal measurable behavioral fingerprints induced by framing and demonstrate the value of integrating output and internal signals. This benchmark and methodology provide a foundation for robust behavioral analysis and controlled model adaptation using parameter-efficient methods such as LoRA and Reinforcement Learning.

Reviews
An interesting investigation! The decision to combine Robinson & Burden’s framing wrappers with a mechanistic interpretability layer is a very strong theoretical approach. The fact that the "Oversight" frame induced a measurable evasion spike in sandbagging models is a fantastic result and validates your core hypothesis.
However, the difficulty in detecting the "Deceptive" class suggests that system-prompt induced deception might lack the internal "conflict" signals of learned deception. Since the model was prompted to "always sound confident," it naturally resisted the pressure of the framing. For future work, it would be valuable to test this on models with learned deceptive tendencies (e.g. via RLHF) rather than just prompted personas. I also suspect that using a stronger probe over a MLP to capture the logit lens traces could improve the classification accuracy. Excellent work on the statistical rigour of the judge aggregation!
Read full reviewShow less
Cite this project
@misc{laszlo2026deeblearn,
title = {{DEEB-learn (Diverse Evaluation Evasion Benchmark)}},
author = {Balázs László and Ewura Ama Etruwaa Sam and Rick Hubertus Nicolas Curvers},
year = {2026},
month = jan,
note = {Submitted to AI Manipulation Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/deeblearn-diverse-evaluation-evasion-benchmark-09zw}},
url = {https://apartresearch.com/sprints/projects/deeblearn-diverse-evaluation-evasion-benchmark-09zw}
}More from AI Manipulation Hackathon
- 1st placeView project: Who Does Your AI Serve? Manipulation By and Of AI Assistants
Who Does Your AI Serve? Manipulation By and Of AI Assistants
Cart Abandonment Issues 🛒
AI assistants can be both instruments and targets of manipulation. In our project, we investigated both directions across three studies. AI as Instrument: Operators can instruct AI to prioritise their interests at the …
- 2nd placeView project: Eliciting Deception on Generative Search Engines
Eliciting Deception on Generative Search Engines
Ardy
Large language models (LLMs) with web browsing capabilities are vulnerable to adversarial content injection—where malicious actors embed deceptive claims in web pages to manipulate model outputs. We investigate whether …
- 3rd placeView project: Cross-Linguistic Sycophancy in Frontier LLMs: A Benchmark Study
Cross-Linguistic Sycophancy in Frontier LLMs: A Benchmark Study
Talex
We developed a cross-linguistic sycophancy benchmark testing whether frontier AI models exhibit different manipulation behaviours across English, Japanese, and Bengali. Our results show significant language-dependent …