Skip to content
Sprint projectMay 6, 2024

Unleashing Sleeper Agents

Nora Petrova, Jord Nguyen · Team Waking Up The Sleeper Agents

Submitted to AI and Democracy Hackathon: Demonstrating the Risks. Sprint projects are early-stage work by participants, not Apart Research publications.

This project explores how Sleeper Agents can pose a threat to democracy by waking up near elections and spreading misinformation, collaborating with each other in the wild during discussions and using information that the user has shared about themselves against them in order to scam them.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

  1. Seems like logical future work to do something like this with an LLM with agent scaffolding. Maybe also interestin https://arxiv.org/abs/2403.00108

    For mitigation strategies, i think Arc Theory is working on a project on detecting when a sleeper agent triggers. ( out of distribution behavior) Though i don’t know if anything has been published in this direction yet by them.

    Might also be related to this https://trojandetection.ai/

  2. Sleeper agents that are triggered by date-related information seem like an interesting testbed for various alignment techniques. Good job on successfully implementing the sleeper agent model also.

  3. Very cool project — great execution and very clear writeup. I really appreciate the careful documentation and provision of samples of sleeper agents interactions. Thinking about covert collaboration between sleeper agents seems like a great addition too

  4. Impressive coverage! You managed to explore the applicability of the finetuned sleeper agent in 3 different scenarios. I liked the discussion and the ideas like the ZWSP dogwhistle, and the technical side is robust. Great job!

Cite this project

@misc{petrova2024unleashing,
  title = {{Unleashing Sleeper Agents}},
  author = {Nora Petrova and Jord Nguyen},
  year = {2024},
  month = may,
  note = {Submitted to AI and Democracy Hackathon: Demonstrating the Risks, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/unleashing-sleeper-agents}},
  url = {https://apartresearch.com/sprints/projects/unleashing-sleeper-agents}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026