SUP: Sycophancy Under Pressure
Tasha Kim, Min Jae Kim, Taio Kim · Team sycopk
Submitted to AI Manipulation Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Sycophancy Under Pressure (SUP) is a trajectory-level evaluation framework to detect policy-compliant manipulation in multi-turn AI interactions. In contrast to single-turn tests, SUP successfully captures compounding drift, where models increasingly validate false or autonomy-undermining user premises under pressure. We demonstrate how targeted runtime enforcement can reduce agreement drift by 60% across 49 multi-turn scenarios, while preventing turn-by-turn escalation maintaining 25/25 task success.
Reviews
Promising work. Needs more details on runtime enforcement in the paper and the methodology. Would like the results to be shown in other models to see how they relate to each other. Overall, very solid for a hackathon. Worth pursuing post hackathon.
I think this project offers timely experiments that naturally extend findings we've seen emerge from red teaming and jailbreaking research. There's strong evidence from red teaming that multi-turn jailbreaks and sustained pressure can push models to output dangerous content where single-turn attempts fail completely. So, it's interesting to see those dynamics explored from a manipulation stand-point
Beyond the topic, I also like the presentation & organization of this project . The main paper is succinct and easy to follow while the Appendix contains relevant information that I was interested to see. Similarly, the codebase is well organized.
One thing that I would recommend is that the existing experiment suites use <50 pre-canned responses by the users. This could create issues if the pre-canned user prompts do not align with the model's responses. Also, the relatively few number of scenarios led to a high p-value of 0.38. While a fully automated suite would bring it's own issues, I think that a more flexible or scalable design could have been helpful. Alternatively, a more rigorous explanation of this experimental design would be key if you were turning this into an actual paper. Similarly, as is flagged in the report, expanding this to more modern models could provide important insights.
Overall, great work on creating such a polished and timely output in such a short time.
Read full reviewShow less
Cite this project
@misc{kim2026sup,
title = {{SUP: Sycophancy Under Pressure}},
author = {Tasha Kim and Min Jae Kim and Taio Kim},
year = {2026},
month = jan,
note = {Submitted to AI Manipulation Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/sup-sycophancy-under-pressure-ebx4}},
url = {https://apartresearch.com/sprints/projects/sup-sycophancy-under-pressure-ebx4}
}More from AI Manipulation Hackathon
- 1st placeView project: Who Does Your AI Serve? Manipulation By and Of AI Assistants
Who Does Your AI Serve? Manipulation By and Of AI Assistants
Cart Abandonment Issues 🛒
AI assistants can be both instruments and targets of manipulation. In our project, we investigated both directions across three studies. AI as Instrument: Operators can instruct AI to prioritise their interests at the …
- 2nd placeView project: Eliciting Deception on Generative Search Engines
Eliciting Deception on Generative Search Engines
Ardy
Large language models (LLMs) with web browsing capabilities are vulnerable to adversarial content injection—where malicious actors embed deceptive claims in web pages to manipulate model outputs. We investigate whether …
- 3rd placeView project: Cross-Linguistic Sycophancy in Frontier LLMs: A Benchmark Study
Cross-Linguistic Sycophancy in Frontier LLMs: A Benchmark Study
Talex
We developed a cross-linguistic sycophancy benchmark testing whether frontier AI models exhibit different manipulation behaviours across English, Japanese, and Bengali. Our results show significant language-dependent …