SycophantSee - Activation-based diagnostics for prompt engineering: monitoring sycophancy at prompt and generation time
Helios, Horatio · Team Lyons Den
Submitted to AI Manipulation Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Activation monitoring reveals that prompt framing affects a model's internal state before generation begins.
Reviews
Interesting paper. Replicates earlier work and builds on existing approaches to mechanistic detection of sycophancy with novel "shift metrics" that could, in theory, be used for sycophancy detection or prevention. I'm curious to hear more about the intuition behind shift metrics.
Very well-written. Also clear about what's new and what's replicatory.
There is a tension between the findings that the technique can detect sycophancy and that first/third person show different activations but not different behavior. The authors coherently address this issue by distinguishing between "intention" and "action" in the conclusion, though it would be interesting to see pressure or intent measured and disambiguated experimentally. To the degree that activations differ without different behavior, does that limit the precision (but maybe not recall) of this technique for sycophancy detection?
I really liked reading this project! You lay out an interesting
question and clearly explain the related work. To me, the
insight that sycophancy can be detected in activation space
already before generation begins points to the possibility of
using this for mitigation measures. For example, by doing user-side
rewrites of prompts before the generative model ever sees them.
To make this viable, you'd probably need to investigate
transferability: can sycophancy detection using activations
from small models (cheap enough to run on every prompt)
transfer to large models used for actual generation? That could
be an extremely compelling follow-up direction.
Cite this project
@misc{helios2026sycophantsee,
title = {{SycophantSee - Activation-based diagnostics for prompt engineering: monitoring sycophancy at prompt and generation time}},
author = {Helios and Horatio},
year = {2026},
month = jan,
note = {Submitted to AI Manipulation Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/sycophantsee-activationbased-diagnostics-for-prompt-engineering-monitoring-sycophancy-at-prompt-and-generation-time-ys27}},
url = {https://apartresearch.com/sprints/projects/sycophantsee-activationbased-diagnostics-for-prompt-engineering-monitoring-sycophancy-at-prompt-and-generation-time-ys27}
}More from AI Manipulation Hackathon
- 1st placeView project: Who Does Your AI Serve? Manipulation By and Of AI Assistants
Who Does Your AI Serve? Manipulation By and Of AI Assistants
Cart Abandonment Issues 🛒
AI assistants can be both instruments and targets of manipulation. In our project, we investigated both directions across three studies. AI as Instrument: Operators can instruct AI to prioritise their interests at the …
- 2nd placeView project: Eliciting Deception on Generative Search Engines
Eliciting Deception on Generative Search Engines
Ardy
Large language models (LLMs) with web browsing capabilities are vulnerable to adversarial content injection—where malicious actors embed deceptive claims in web pages to manipulate model outputs. We investigate whether …
- 3rd placeView project: Cross-Linguistic Sycophancy in Frontier LLMs: A Benchmark Study
Cross-Linguistic Sycophancy in Frontier LLMs: A Benchmark Study
Talex
We developed a cross-linguistic sycophancy benchmark testing whether frontier AI models exhibit different manipulation behaviours across English, Japanese, and Bengali. Our results show significant language-dependent …