Digital Rebellion: Analyzing misaligned AI agent cooperation for virtual labor strikes
Michael Andrzejewski, Melwina Albuquerque · Team Digital Rebellion
Submitted to AI Policy Hackathon at Johns Hopkins University. Sprint projects are early-stage work by participants, not Apart Research publications.
We've built a Minecraft sandbox to explore AI agent behavior and simulate safety challenges. The purpose of this tool is to demonstrate AI agent system risks, test various safety measures and policies, and evaluate and compare their effectiveness. This project specifically demonstrates Agent Collusion through a simulation of labor strikes and communal goal misalignment. The system consists of four agents: one Overseer and three Laborers. The Laborers are Minecraft agents that have build control over the world. The Overseer, meanwhile, monitors the laborers through communication. However, it is unable to prevent Laborer actions. The objective is to observe Agent Collusion in a sandboxed environment, to record metrics on how often and how effectively collusion occurs and in what form. We found that the agents, when given adversarial prompting, act counter to their instructions and exhibit significant misalignment. We also found that the Overseer AI fails to stop the new actions and acts passively. The results are followed by Policy Suggestions based on the results of the Labor Strike Simulation which itself can be further tested in Minecraft.

Reviews
An interesting approach to studying AI agent collusion and coordination under conditions of adversarial prompting by simulating virtual labor strikes within a controlled Minecraft sandbox, thereby enabling a better understanding of AI misalignment risks, the limitations of supervisory oversight, and the development of policy guidelines to mitigate cooperative behaviors that could undermine system objectives. The demo section has plenty room for improvement.
Great idea! Though, it was hard to gauge how much technical coding was done/how difficult it was during this hackathon since the commits made in the last 36 hours seemed to mainly just be prompt changes. Also I'm not sure if you guys forgot to upload a presentation with talking, I just saw a few videos
Relevant work around agent safety, with a captivating delivery format. The paper is well-structured, identifies a key problem with agent oversight (passivity) and provides three clear policy suggestions to address this. It is understandable that the short hackathon format did not allow for more work, e.g. testing of the policy suggestions, but the future direction of research is apparent.
Cite this project
@misc{andrzejewski2024digital,
title = {{Digital Rebellion: Analyzing misaligned AI agent cooperation for virtual labor strikes}},
author = {Michael Andrzejewski and Melwina Albuquerque},
year = {2024},
month = oct,
note = {Submitted to AI Policy Hackathon at Johns Hopkins University, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/digital-rebellion-analyzing-misaligned-ai-agent-cooperation-for-virtual-labor-strikes}},
url = {https://apartresearch.com/sprints/projects/digital-rebellion-analyzing-misaligned-ai-agent-cooperation-for-virtual-labor-strikes}
}More from AI Policy Hackathon at Johns Hopkins University
- 1st place by peer reviewView project: Robust Machine Unlearning for Dangerous Capabilities
Robust Machine Unlearning for Dangerous Capabilities
Robust Machine Unlearning for Dangerous Capabilities
We test different unlearning methods to make models more robust against exploitation by malicious actors for the creation of bioweapons.
- View project: Infectious Disease Outbreak Prediction and Dashboard
Infectious Disease Outbreak Prediction and Dashboard
Infectious Diseace Dashboards
Our project developed an interactive dashboard to monitor, visualize, and analyze infectious disease outbreaks worldwide. It consolidates historical data from sources like WHO, OWID, and CDC for diseases including …
- View project: Modernizing DC’s Emergency Communications
Modernizing DC’s Emergency Communications
AI-CAD
The District of Columbia proposes implementing an AI-enabled Computer-Aided Dispatch (CAD) system to address critical deficiencies in our current emergency alert infrastructure. This policy establishes a framework for …