Skip to content
Sprint projectOct 6, 2024

AI Honeypot

Reworr

Submitted to Agent Security Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

The project designed to monitor AI Hacking Agents in the real world using honeypots with prompt injections and temporal analysis.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

  1. A creative approach to using honeypots with prompt injection techniques is both fascinating and effective. The project’s ability to gather real-world data on AI behavior is a game-changer, offering a unique glimpse into emerging cybersecurity challenges. It’s a true standout in the field of agent safety!

  2. Great deliverables. The website containing the explainer and paper together with the animated slides improve the communication aspect significantly.• The idea temporal analysis to distinguish LLM agents from humans based on response times is particularly good.

  3. I find your project to be an innovative and timely approach to tackling AI-driven cybersecurity threats. Your integration of real-time metrics through a public-facing dashboard is impressive, offering transparency and practical insights to the cybersecurity community. One area for improvement could be around increasing the dataset of AI-driven hacking attempts, which seems currently limited. This could help in refining detection methods further and provide more robust insights.

  4. This is a great idea on securing systems from AI agents’s attacks. I would love to see two things in future releases: (a) False positive and False negative rates of Honeypot, (b): I believe LLM agents are becoming quite smart w.r.t. prompt injections. As such, do the prompt injection methods  work over time?

Cite this project

@misc{reworr2024ai,
  title = {{AI Honeypot}},
  author = {Reworr},
  year = {2024},
  month = oct,
  note = {Submitted to Agent Security Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/ai-honeypot}},
  url = {https://apartresearch.com/sprints/projects/ai-honeypot}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026