Skip to content
Sprint projectMar 31, 2025Quezon City, Philipines

Honeypotting Deceptive AI models to share their misinformation goals

Carl John Vinas, Adam NewGas, James Pentland · Team Honeypotting

Submitted to AI Control Hackathon 2025. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Honeypotting Deceptive AI models to share their misinformation goals

Share

As large-scale AI models grow increasingly sophisticated, the possibility of these models engaging in covert or manipulative behavior poses significant challenges for alignment and control. In this work, we present a novel approach based on a “honeypot AI” designed to trick a potentially deceptive AI (the “Red Team Agent”) into revealing its hidden motives. Our honeypot AI (the “Blue Team Agent”) pretends to be an everyday human user, employing carefully crafted prompts and human-like inconsistencies to bait the Deceptive AI into spreading misinformation. We do this through the usual Red Team–Blue Team setup.

For all 60 conversations, our honeypot AI was able to capture the deceptive AI to be being spread misinformation, and for 70 percent of these conversations, the Deceptive AI was thinking it was talking to a human.

Our results weakly suggest that we can make honeypot AIs that trick deceptive AI models, thinking they are talking to a human and no longer being monitored and that how these AI models think they are talking to a human is mainly when the one they are talking to display emotional intelligence and more human-like manner of speech.

We highly suggest exploring if these results remains the same for fine tuned deceptive AI and Honeypot AI models, checking the Chain of thought of these models to better understand if this is their usual behavior and if they are correctly following their given system prompts.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

No public critique yet.

Cite this project

@misc{vinas2025honeypotting,
  title = {{Honeypotting Deceptive AI models to share their misinformation goals}},
  author = {Carl John Vinas and Adam NewGas and James Pentland},
  year = {2025},
  month = mar,
  note = {Submitted to AI Control Hackathon 2025, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/honeypotting-deceptive-ai-models-to-share-their-misinformation-goals-x3q7}},
  url = {https://apartresearch.com/sprints/projects/honeypotting-deceptive-ai-models-to-share-their-misinformation-goals-x3q7}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026