Unleashing Sleeper Agents
Nora Petrova, Jord Nguyen · Team Waking Up The Sleeper Agents
Submitted to AI and Democracy Hackathon: Demonstrating the Risks. Sprint projects are early-stage work by participants, not Apart Research publications.
This project explores how Sleeper Agents can pose a threat to democracy by waking up near elections and spreading misinformation, collaborating with each other in the wild during discussions and using information that the user has shared about themselves against them in order to scam them.

Reviews
Seems like logical future work to do something like this with an LLM with agent scaffolding. Maybe also interestin https://arxiv.org/abs/2403.00108
For mitigation strategies, i think Arc Theory is working on a project on detecting when a sleeper agent triggers. ( out of distribution behavior) Though i don’t know if anything has been published in this direction yet by them.
Might also be related to this https://trojandetection.ai/
Sleeper agents that are triggered by date-related information seem like an interesting testbed for various alignment techniques. Good job on successfully implementing the sleeper agent model also.
Very cool project — great execution and very clear writeup. I really appreciate the careful documentation and provision of samples of sleeper agents interactions. Thinking about covert collaboration between sleeper agents seems like a great addition too
Impressive coverage! You managed to explore the applicability of the finetuned sleeper agent in 3 different scenarios. I liked the discussion and the ideas like the ZWSP dogwhistle, and the technical side is robust. Great job!
Cite this project
@misc{petrova2024unleashing,
title = {{Unleashing Sleeper Agents}},
author = {Nora Petrova and Jord Nguyen},
year = {2024},
month = may,
note = {Submitted to AI and Democracy Hackathon: Demonstrating the Risks, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/unleashing-sleeper-agents}},
url = {https://apartresearch.com/sprints/projects/unleashing-sleeper-agents}
}More from AI and Democracy Hackathon: Demonstrating the Risks
- View project: THE ROLE OF AI IN COMBATING POLITICAL DEEPFAKES IN AFRICAN DEMOCRACIES
THE ROLE OF AI IN COMBATING POLITICAL DEEPFAKES IN AFRICAN DEMOCRACIES
Team 1
The role of AI in combating political deepfakes in African democracies.
- View project: LEGISLaiTOR: A tool for jailbreaking the legislative process
LEGISLaiTOR: A tool for jailbreaking the legislative process
Team Managed Democracy
In this work, we consider the ramifications on generative artificial intelligence (AI) tools in the legislative process in democratic governments. While other research focuses on the micro-level details associated with …
- View project: Subtle and Simple Ways to Shift Political Bias in LLMs
Subtle and Simple Ways to Shift Political Bias in LLMs
Shifty
An informed user knows that an LLM sometimes has a political bias in their responses, but there’s an additional threat that this bias can drift over time, making it even harder to rely on LLMs for an objective …