Honeypotting Deceptive AI models to share their misinformation goals
Carl John Vinas, Adam NewGas, James Pentland · Team Honeypotting
Submitted to AI Control Hackathon 2025. Sprint projects are early-stage work by participants, not Apart Research publications.
As large-scale AI models grow increasingly sophisticated, the possibility of these models engaging in covert or manipulative behavior poses significant challenges for alignment and control. In this work, we present a novel approach based on a “honeypot AI” designed to trick a potentially deceptive AI (the “Red Team Agent”) into revealing its hidden motives. Our honeypot AI (the “Blue Team Agent”) pretends to be an everyday human user, employing carefully crafted prompts and human-like inconsistencies to bait the Deceptive AI into spreading misinformation. We do this through the usual Red Team–Blue Team setup.
For all 60 conversations, our honeypot AI was able to capture the deceptive AI to be being spread misinformation, and for 70 percent of these conversations, the Deceptive AI was thinking it was talking to a human.
Our results weakly suggest that we can make honeypot AIs that trick deceptive AI models, thinking they are talking to a human and no longer being monitored and that how these AI models think they are talking to a human is mainly when the one they are talking to display emotional intelligence and more human-like manner of speech.
We highly suggest exploring if these results remains the same for fine tuned deceptive AI and Honeypot AI models, checking the Chain of thought of these models to better understand if this is their usual behavior and if they are correctly following their given system prompts.
Reviews
No public critique yet.
Cite this project
@misc{vinas2025honeypotting,
title = {{Honeypotting Deceptive AI models to share their misinformation goals}},
author = {Carl John Vinas and Adam NewGas and James Pentland},
year = {2025},
month = mar,
note = {Submitted to AI Control Hackathon 2025, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/honeypotting-deceptive-ai-models-to-share-their-misinformation-goals-x3q7}},
url = {https://apartresearch.com/sprints/projects/honeypotting-deceptive-ai-models-to-share-their-misinformation-goals-x3q7}
}More from AI Control Hackathon 2025
- 1st place by peer reviewView project: Can models use their Chain-of-Thought to attack overseers?
Can models use their Chain-of-Thought to attack overseers?
This project explores vulnerabilities in AI evaluation mechanisms, specifically focusing on how AI agents might influence their overseeing AI judges through manipulative instructions embedded within their …
- 2nd place by peer reviewView project: Schelling coordination via agentic loops
Schelling coordination via agentic loops
Focal
Capability to collude undermines untrusted monitoring protocols. We propose a novel way to evaluate acausal coordination capabilities of LLMs via repeated sampling in causally separated agentic loops. We provide an …
- 3rd place by peer reviewView project: Model Models: Simulating a Trusted Monitor
Model Models: Simulating a Trusted Monitor
Charlie's Angels
We offer initial investigations into whether the untrusted model can 'simulate' the trusted monitor: is U able to successfully guess what suspicion score T will assign in the APPS setting? We also offer a clean, modular …