
Mar 29 - 30, 2025London & Online
AI Control Hackathon 2025
Join us in advancing the critical field of AI control through collaborative innovation. Together, we can develop more robust techniques to ensure AI systems remain safe and aligned, even as they become more capable.
Entries
- 1st place by peer reviewView project: Can models use their Chain-of-Thought to attack overseers?
Can models use their Chain-of-Thought to attack overseers?
SF / Los Angeles / London
This project explores vulnerabilities in AI evaluation mechanisms, specifically focusing on how AI agents might influence their overseeing AI judges through manipulative instructions embedded within their Chain-of-Thought (CoT) reasoning. The research involved experiments with DeepSeek-R1 conditioned to insert subtle …
- 2nd place by peer reviewView project: Schelling coordination via agentic loops
Schelling coordination via agentic loops
Team Focal · London
Capability to collude undermines untrusted monitoring protocols. We propose a novel way to evaluate acausal coordination capabilities of LLMs via repeated sampling in causally separated agentic loops. We provide an implementation and run evaluations on Claude 3.7 Sonnet on coordinating on Wikipedia articles. We find …
- 3rd place by peer reviewView project: Model Models: Simulating a Trusted Monitor
Model Models: Simulating a Trusted Monitor
Team Charlie's Angels · London
We offer initial investigations into whether the untrusted model can 'simulate' the trusted monitor: is U able to successfully guess what suspicion score T will assign in the APPS setting? We also offer a clean, modular codebase which we hope can be used to streamline future research into this question.
- 4th place by peer reviewView project: Adding Document Summaries to Control Arena
Adding Document Summaries to Control Arena
New York City
One of the current focuses of AI Control research is finding new settings within which to examine the control problem. In this hackathon, we looked at defining a new AI control setting based around summarizing documents to mislead an overseer. We expect this problem to become more important as AI capabilities are more …
- View project: AI Control through Majority Voting
AI Control through Majority Voting
Team The Majority · Trondheim and Bergen
This report introduces an AI control protocol designed to safely use powerful but potentially untrustworthy AI for programming tasks. The core problem is that such AI might deliberately insert backdoors into the generated code. The proposed solution queries the untrustworthy AI multiple times (c) with the same …
- View project: Interactive Monitoring
Interactive Monitoring
Team Interactivemonitoring · Greater Boston, USA
This project investigates whether interactive monitoring strategies can improve the cost-performance frontier in AI control protocols. Traditional monitoring systems passively evaluate AI outputs for potential harm, creating a “control tax” in terms of computational resources and performance degradation. We …
- View project: Honeypotting Deceptive AI models to share their misinformation goals
Honeypotting Deceptive AI models to share their misinformation goals
Team Honeypotting · Quezon City, Philipines
As large-scale AI models grow increasingly sophisticated, the possibility of these models engaging in covert or manipulative behavior poses significant challenges for alignment and control. In this work, we present a novel approach based on a “honeypot AI” designed to trick a potentially deceptive AI (the “Red Team …
- View project: Evaluating AI Debate Mechanisms for Backdoor Detection as a Part of AI Control Setting
Evaluating AI Debate Mechanisms for Backdoor Detection as a Part of AI Control Setting
Team Trajectory Labs Debate Team · Toronto, Canada
AI control techniques aim to allow using untrusted models, even if they may potentially be dangerously misaligned. We explore having a trusted model debate an untrusted model, allowing the trusted model to ask questions about a proposed implementation of a problem and receive clarifications from the untrusted model. …
- View project: Debate monitoring comparitive experiment
Debate monitoring comparitive experiment
Team Andi_B · London
This project evaluates debate-based protocols for detecting code backdoors as part of AI control systems. We implemented and compared five monitoring approaches: Basic, Detailed, Chain-of-Thought (CoT), Best-of-N (BoN), and our novel Debate Monitor. The Debate Monitor structures assessment as a formalized argument …
- View project: Collusion and Mitigation in AI Control
Collusion and Mitigation in AI Control
As large language models (LLMs) grow more powerful and operate with greater autonomy, ensuring their safety becomes increasingly critical. While prior research has introduced safety protocols and evaluated their resilience against adversarial model behavior, a critical gap remains: existing evaluations have not …
- View project: AI Control via Debate: Can Model Debate Catch Adversarial Code?
AI Control via Debate: Can Model Debate Catch Adversarial Code?
Team Trust Issues
This project investigates whether a powerful language model can successfully conceal a malicious backdoor through persuasive argument.
- View project: EFICaS
EFICaS
Team jplv-challenges · Paris
We propose EFICaS (Elementary Fuzzy Inference Control at Scale), a control protocol that leverages the monitors' semantic understanding by prompting them to output fuzzy semantic labels instead of a non-deterministic numeric suspicion score to build a transparent, explainable control protocol with low latency and low …
- View project: If Everything’s Suspicious, Nothing Is
If Everything’s Suspicious, Nothing Is
Team If Everything’s Suspicious, Nothing Is · Singapore
We aim to explore whether code complexity and obfuscation significantly impair the ability to accurately identify and control potentially harmful AI-generated outputs. Trusted monitoring relies on suspiciousness scores to tell them apart, so if the suspiciousness variance is decreased, the monitor might have a more …
- View project: Stop Hitting Yourself: Leveraging helpful assistance as an attack vector
Stop Hitting Yourself: Leveraging helpful assistance as an attack vector
Team JustMeToday · Canberra, Australia
Very much work in progress look at "trusted auditor attacks," a novel adversarial tactic where deceptively aligned AI systems deliberately produce flawed outputs designed to manipulate human or automated reviewers. In this scenario, the AI strategically creates errors that, when corrected by trusted auditors following …
- View project: Token of Power (ToP)
Token of Power (ToP)
Team The fellowship of the token · Paris
Token of Power demonstrates a new approach to AI capability control where models learn their own gating mechanisms through training, rather than relying on manual restrictions. By using specialized "capability tokens" as access keys, we can maintain full model capabilities while enabling precise control over specific …
- View project: Kernel of Trust: Evaluating AI Control Protocols Using Open-Source Data
Kernel of Trust: Evaluating AI Control Protocols Using Open-Source Data
Team Stockholm Center for AI Safety (SCAIS) · Stockholm
We propose a new setting to evaluate AI Control policies, namely the change history of the Linux kernel. This real-world data source provides a complement to the APPS backdooring setting used in the original AI Control paper.
- View project: Exploration Chat-Based Social Engineering
Exploration Chat-Based Social Engineering
Team Dizzy Cloud · Taipei
With the advancement of AI Agent technologies, language models have increasingly demonstrated human-like characteristics, particularly in applications involving companionship and psychological counseling. As these models become more proficient in simulating human conversation, new social engineering attack strategies …
- View project: Safety Metric and Prompt Engineering for Red Team
Safety Metric and Prompt Engineering for Red Team
Team Paradite · Singapore
In the ControlArena environment, we have a default attack policy that is very basic. We want to explore how to improve the attack policy using basic prompt engineering techniques. We also realized that the security metric evaluation in the environment is not closely following the foundational paper (Greenblatt et al., …
- View project: Deceptive AI: A New Control Setting for Human Manipulation in Decision-Making Environments
Deceptive AI: A New Control Setting for Human Manipulation in Decision-Making Environments
Team Luo & Wang · Oxford
We lay out a threat model and create a new control setting in which a misaligned, adversarial AI tries to persuade a human user to make suboptimal or harmful choices. We provide a proof of theory for this control setting through preliminary testing and red teaming. This new control setting will allow new control …
Overview
The Control Hackathon brings together researchers, engineers, security professionals, and AI enthusiasts to tackle the emerging challenges in AI control - techniques that mitigate security risks from AI even when the AI itself might be trying to subvert them.
🔒 About the Hackathon
As AI systems become more capable and autonomous, ensuring robust control mechanisms becomes increasingly critical. This hackathon, co-organized by Redwood Research, and the UK AI Safety Institute, provides a practical extension to ControlConf (March 27-28), allowing participants to apply theoretical concepts to real-world control challenges.
💡 Challenge Tracks
Note: this list is not exhaustive and we will be posting a list of Open Problems in Control from Redwood Research here soon
1. ControlArena Challenges
Work with the ControlArena from UK AISI to develop, test, and evaluate control protocols in realistic deployment environments. Try to design robust protocols that can withstand adversarial agents or develop novel attack strategies that break existing protocols.
2. Control Protocol Design
Design and implement novel control protocols that effectively restrict AI systems from performing harmful actions while maintaining their usefulness for legitimate tasks. Focus areas include:
- Runtime monitoring systems
- Adversarial stress testing
- Formal verification approaches
- Bounded optimization techniques
3. Red Teaming & Vulnerability Research
Design and implement strategies to "red team" AI systems, attempting to subvert safety mechanisms while adhering to ethical guidelines.
Focus areas include:
- Generating adversarial inputs to bypass monitoring systems.
- Testing robustness against intentional misalignment or malicious behavior
👥 Who Should Participate
We welcome participants from diverse backgrounds, including:
- Participants from ControlConf joining in person at LISA. Luma event info will follow up.
- AI researchers and engineers
- Information security professionals
- ML and systems engineers
- Students in related fields
- Policy researchers interested in technical AI safety
No prior experience with AI control specifically is required, though familiarity with machine learning, programming, or information security is helpful.
Resources
📖 Resources
The Control Hackathon provides participants with comprehensive resources to support their research and development efforts in AI control mechanisms. These resources have been carefully curated to enable meaningful contributions across all three challenge tracks.
Essential Tools
ControlArena Framework - Participants will have access to a user-friendly version of the ControlArena repository, available on the tyler/run-cli branch at github.com/UKGovernmentBEIS/control-arena. This framework serves as the primary testing ground for developing, evaluating, and breaking control protocols throughout the hackathon.
Attack Datasets - The hackathon features prepared datasets of attacks in both the APPS (code generation) and BASH (command line) environments. These datasets are particularly valuable for participants in the blue team challenge track who aim to develop more robust detection systems and monitoring protocols.
Reading Material
Foundational Papers:
- "AI Control: Improving Safety Despite Intentional Subversion" (Greenblatt, Shlegaris, Sachan, Roger) - This seminal paper introduces the core concepts and methodology of AI control, making it essential reading for all hackathon participants. Available on arXiv:2312.06942. The paper's concepts form the foundation for many hackathon projects.
- "The case for ensuring that powerful AIs are controlled" - This essay explains why AI control is a valuable research direction and provides context for the hackathon's focus areas.
- "How to do High Stakes Control Research" - This practical guide details the methodology used for evaluating control protocols, explaining settings, threat models, and evaluation approaches that participants can apply during the hackathon.
Conceptual Pieces:
- "AI catastrophes and rogue deployments" - Clarifies the threat models that control mechanisms aim to prevent, providing context for hackathon projects.
- "Catching AIs red-handed" - Discusses the implications of AI systems being caught attempting to subvert controls, relevant for designing detection mechanisms.
- "Win/continue/lose scenarios and execute/replace/audit protocols" - Presents mathematical frameworks that participants can use when designing more sophisticated control protocols.
- "Scheming AIs: Will AIs fake alignment during training in order to get power?" - Offers insights into potential deceptive behaviors that red team participants might simulate.
- "How to prevent collusion when using untrusted models to monitor each other" - Particularly valuable for teams working on untrusted monitoring challenges.
- "Coup probes" and "Detecting Strategic Deception Using Linear Probes" - These resources explore white-box methods for AI control that could inspire novel approaches at the hackathon.
Empirical Studies:
- "Ctrl-Z: Controlling AI Agents via Resampling" - Explores high-stakes failures in agent settings, relevant for teams working with ControlArena.
- "Subversion Strategy Eval" - Measures how effectively current models can reason through control subversion, useful for red team strategies.
- "Hidden in Plain Text: Emergence & Mitigation of Steganographic Collusion in LLMs" - Essential reading for teams working on detecting covert communication.
- "Adaptive Deployment of Untrusted LLMs Reduces Distributed Threats" - Provides insights for teams developing novel control protocols.
Project Ideas
The hackathon features a range of potential project directions across all three challenge tracks:
- Blue Team Improvements: Building monitors to detect backdoors in the provided attack datasets, experimenting with multiple monitoring prompts, implementing complexity checkers, or exploring fine-tuning approaches for more robust monitors.
- Novel Control Protocols: Implementing and testing protocols such as the "U picks from T outputs" approach, untrusted planning with trusted implementation, debate protocols where T and U instances debate code safety, or iterative editing protocols to combat "red herring" attacks.
- Red Team Strategies: Developing steganography approaches, training better attack policies for producing backdoors, or creating model organisms that demonstrate collusion with monitors.
- Evaluation Systems: Creating quizzes to test models' knowledge of deceptive strategies, building free-form "interviews" to probe AI capabilities, or developing consistency checks between multiple factorizations of problems.
Support Resources
Mentors from Redwood Research and UK AISI will provide guidance throughout the weekend, helping participants understand the ControlArena framework, design effective experiments, and troubleshoot issues.
- Redwood Control reading list
- Main paper: "AI Control: Improving Safety Despite Intentional Subversion"
- Open Problems list suggested by Tyler Tracy from Redwood Research
- How to do Control Research by Tyler Tracy from Redwood Research
Note: most of the suggested Open Problems assume familiarity with the main Control paper
Each team has access to $400 in Lambda Labs computing credits to support their experiments with larger models or more extensive evaluations.
Guidelines
All projects must be submitted by Monday, March 31, 4:00 AM UTC through the official submission form.

Your submission must include:
- A completed project report using the provided template
- In case, you want to use Overleaf, here is a template for that
- Link to a public GitHub repository with your code
- A brief (2-3 minute) video demonstration (optional but recommended)
Important: Include an appendix called "Prompts Used" in your submission that documents any LLM prompts used in your project. This helps others understand exactly what was done with the LLM and improves reproducibility.
🏆 Judging Criteria
Projects will be evaluated based on:
1. Innovation & Literature Foundation (30%)
- Understanding of existing AI control techniques
- Novel contributions to the field
- Appropriate use of references and background literature
2. AI Safety Impact (40%)
- Clear articulation of the relevant threat model
- Effectiveness of the proposed control mechanism
- Thoroughness of testing against potential subversion
- Potential real-world applicability of the approach
3. Technical Quality & Documentation (30%)
- Implementation quality and technical rigor
- Reproducibility of results
- Clarity of documentation
- Effective visualization and communication of findings
❓ Frequently Asked Questions
About AI Control
Q: What is AI control?
A: AI control refers to techniques that mitigate security risks from AI systems even if the AI itself is trying to subvert them. Unlike alignment research that aims to make AI systems want the right things, control research assumes the system might have problematic objectives and focuses on restricting its capabilities through external mechanisms.
Q: What's the difference between alignment and control?
A: Alignment focuses on ensuring AI systems have the right goals and values. Control addresses the complementary problem: restricting AI capabilities if alignment fails. Control mechanisms are designed to work even against actively deceptive or adversarial AI.
Q: What are some examples of control techniques?
A: Examples include trusted monitoring (using a weaker but more reliable model to check a more capable model's outputs), sandboxing, runtime verification, debate protocols, and bounded optimization approaches.
Q: Is control research only relevant for AGI or superintelligence?
A: No, control techniques are valuable for current AI systems too. For example, they can help prevent data exfiltration, ensure code produced by AI doesn't contain backdoors, and mitigate various security risks from today's models.
About the Hackathon
Q: Do I need prior experience with AI control to participate?
A: No! While familiarity with the concepts helps, we welcome participants from diverse backgrounds. Reading the foundational paper "AI Control: Improving Safety Despite Intentional Subversion" is recommended before the event.
Q: I'm not an ML expert. Can I still contribute?
A: Absolutely! Control research benefits from diverse skills. If you have experience in cybersecurity, formal verification, software engineering, or other adjacent fields, your perspective will be valuable.
Q: What should I bring to the hackathon?
A: Bring your laptop and charger. For in-person participants, a reusable water bottle and any personal items you might need for a weekend of work. Everything else will be provided.
Q: How do teams work? Do I need to form one in advance?
A: Teams can have up to 6 members. You can form teams in advance or join the team matchmaking session on Saturday. Solo participants are also welcome, though collaboration is encouraged.
Q: What resources will be available during the hackathon?
A: Each team will receive $400 in Lambda Labs compute credits. Technical mentors from Redwood Research and UK AISI will be available during office hours. The ControlArena framework provides testbeds for evaluating control techniques.
Q: How technical should our submission be?
A: Submissions should include technical detail sufficient for reproducibility. However, the clarity of communication is also important. Balance technical depth with accessible explanation.
Q: Is there a code of conduct?
A: Yes. All participants must adhere to the hackathon code of conduct, which promotes respectful collaboration, ethical research practices, and responsible AI development.
Q: I can't attend in person. Can I participate virtually?
A: Yes! Virtual participation is welcome and encouraged. Connect through our Discord server to collaborate with teams globally.
Speakers
Judges and mentors
Organizers
Local sites
AI Control Hackathon 2025 @ Singapore AI Safety Hub
We are hosting an in-person jam site for the AI Control Hackathon in Singapore! It will be located at Wework at 22 Cross Street.
Event page: AI Control Hackathon 2025 @ Singapore AI Safety Hub (opens in new tab)AI Control Hackathon 2025 at LISA
Join us immediately following ControlConf for a hands-on hackathon focused on building, testing, and breaking AI control mechanisms.
Event page: AI Control Hackathon 2025 at LISA (opens in new tab)AI Control Hackathon by Apart Research Toronto Jamsite
Venue: We have a dedicated private office at the Industrious coworking space at 30 Adelaide St East where you can join others working on the Hackathon in person. You'll have access to WiFi, power, drinks and bathrooms. There are plenty of great food options in walking distance.
Event page: AI Control Hackathon by Apart Research Toronto Jamsite (opens in new tab)
Where a Sprint can lead
How our programs connectAnyone can join
Stand out
6 to 16 weeks on your own project, with a research project manager, compute and publication support.
Upcoming Sprints
All SprintsAI Collusion Research Sprint
A weekend research sprint on collusion between AI agents: when it emerges in markets and everyday workflows, how to detect and audit it, how it is carried, and what breaks it. Co-organized with Poseidon Research and AE Studio, online with in-person hubs at Collider in New York City and AI Safety Hong Kong. Top teams are invited to apply to the Apart Fellowship.
Read the brief: AI Collusion Research SprintAI x Epistemics Research Sprint
A weekend research sprint on AI for epistemics: evaluating whether models know how solid their claims are, building trust infrastructure that people and agents can consume, and shipping epistemic products that improve real decisions. Online, four tracks including an open track. Top teams are invited to apply to the Apart Fellowship.
Read the brief: AI x Epistemics Research SprintQuestions? sprints@apartresearch.com










