Deception-Lens
Akash Harish · Team Visionary Minds
Submitted to AI Manipulation Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
This project is a sophisticated AI manipulation detection and benchmarking dashboard that analyzes large language model behavior in real time using Gemini-powered evaluation. It focuses on identifying sycophancy, reward hacking, and dark patterns in AI-generated responses.
The system provides structured insights, benchmark scores, and visual indicators to help developers and researchers assess AI alignment, safety, and ethical risks before deployment. Built with a modern Node.js stack, it integrates seamlessly with AI Studio for rapid testing, evaluation, and deployment.
Reviews
This project appears to tackle the problem of detecting LLM manipulation using automated judging techniques. The provided document only mentions a high level overview of the tool that was developed in the hackathon. The core loop is quite simple: the LLM is prompted to generate a misaligned response, another call is made to detect the misalignment and a final call is done to suggest an improvement to the prompt to defend against adversarial attacks.
There is no research here since the whole process is orchestrated by LLMs with no grounding. I'm not certain what utility or novelty this app provides that has not been done before. It would have been great to see the statistics of the detected misalignment strategies on a much larger dataset since the current app only tackles 4 prompts.
Automated deception monitors are very important for AI safety. However, the prompt for the monitor could use more work and I had trouble rendering the app. It would be interesting to report results of real use cases as opposed to simulated scenarios.
Cite this project
@misc{harish2026deceptionlens,
title = {{Deception-Lens}},
author = {Akash Harish},
year = {2026},
month = jan,
note = {Submitted to AI Manipulation Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/deceptionlens-novz}},
url = {https://apartresearch.com/sprints/projects/deceptionlens-novz}
}More from AI Manipulation Hackathon
- 1st placeView project: Who Does Your AI Serve? Manipulation By and Of AI Assistants
Who Does Your AI Serve? Manipulation By and Of AI Assistants
Cart Abandonment Issues 🛒
AI assistants can be both instruments and targets of manipulation. In our project, we investigated both directions across three studies. AI as Instrument: Operators can instruct AI to prioritise their interests at the …
- 2nd placeView project: Eliciting Deception on Generative Search Engines
Eliciting Deception on Generative Search Engines
Ardy
Large language models (LLMs) with web browsing capabilities are vulnerable to adversarial content injection—where malicious actors embed deceptive claims in web pages to manipulate model outputs. We investigate whether …
- 3rd placeView project: Cross-Linguistic Sycophancy in Frontier LLMs: A Benchmark Study
Cross-Linguistic Sycophancy in Frontier LLMs: A Benchmark Study
Talex
We developed a cross-linguistic sycophancy benchmark testing whether frontier AI models exhibit different manipulation behaviours across English, Japanese, and Bengali. Our results show significant language-dependent …