Skip to content
AI Control Hackathon 2025

Mar 29 - 30, 2025London & Online

AI Control Hackathon 2025

This Sprint has ended.

Sign-ups
182
Projects submitted
19
Browse the 19 projects

Submit your project

Project details

A short abstract: what you did, what you found.

PDF, up to 25 MB.

Are you interested in publishing this project? *
Tracks

If this Sprint has numbered tracks, choose the ones your project fits.

PDF, PowerPoint, Keynote or ODP, up to 25 MB.

PNG, JPEG, WebP or GIF, up to 25 MB.

Team details

Team member 1

Leave blank if you don’t have one.

By submitting you agree to the prize terms and our Privacy Policy.

The submission window has closed. Your draft is still here so you can copy it, but it can no longer be submitted.

See upcoming Sprints

Join us in advancing the critical field of AI control through collaborative innovation. Together, we can develop more robust techniques to ensure AI systems remain safe and aligned, even as they become more capable.

Entries

Overview

The Control Hackathon brings together researchers, engineers, security professionals, and AI enthusiasts to tackle the emerging challenges in AI control - techniques that mitigate security risks from AI even when the AI itself might be trying to subvert them.

🔒 About the Hackathon

As AI systems become more capable and autonomous, ensuring robust control mechanisms becomes increasingly critical. This hackathon, co-organized by Redwood Research, and the UK AI Safety Institute, provides a practical extension to ControlConf (March 27-28), allowing participants to apply theoretical concepts to real-world control challenges.

💡 Challenge Tracks

Note: this list is not exhaustive and we will be posting a list of Open Problems in Control from Redwood Research here soon

1. ControlArena Challenges

Work with the ControlArena from UK AISI to develop, test, and evaluate control protocols in realistic deployment environments. Try to design robust protocols that can withstand adversarial agents or develop novel attack strategies that break existing protocols.

2. Control Protocol Design

Design and implement novel control protocols that effectively restrict AI systems from performing harmful actions while maintaining their usefulness for legitimate tasks. Focus areas include:

  • Runtime monitoring systems
  • Adversarial stress testing
  • Formal verification approaches
  • Bounded optimization techniques

3. Red Teaming & Vulnerability Research

Design and implement strategies to "red team" AI systems, attempting to subvert safety mechanisms while adhering to ethical guidelines.

Focus areas include:

  • Generating adversarial inputs to bypass monitoring systems.
  • Testing robustness against intentional misalignment or malicious behavior

‍

👥 Who Should Participate

We welcome participants from diverse backgrounds, including:

  • Participants from ControlConf joining in person at LISA. Luma event info will follow up.
  • AI researchers and engineers
  • Information security professionals
  • ML and systems engineers
  • Students in related fields
  • Policy researchers interested in technical AI safety

No prior experience with AI control specifically is required, though familiarity with machine learning, programming, or information security is helpful.

‍

Resources

📖 Resources

The Control Hackathon provides participants with comprehensive resources to support their research and development efforts in AI control mechanisms. These resources have been carefully curated to enable meaningful contributions across all three challenge tracks.

Essential Tools

ControlArena Framework - Participants will have access to a user-friendly version of the ControlArena repository, available on the tyler/run-cli branch at github.com/UKGovernmentBEIS/control-arena. This framework serves as the primary testing ground for developing, evaluating, and breaking control protocols throughout the hackathon.

Attack Datasets - The hackathon features prepared datasets of attacks in both the APPS (code generation) and BASH (command line) environments. These datasets are particularly valuable for participants in the blue team challenge track who aim to develop more robust detection systems and monitoring protocols.

Reading Material

Foundational Papers:

Conceptual Pieces:

Empirical Studies:

Project Ideas

The hackathon features a range of potential project directions across all three challenge tracks:

  • Blue Team Improvements: Building monitors to detect backdoors in the provided attack datasets, experimenting with multiple monitoring prompts, implementing complexity checkers, or exploring fine-tuning approaches for more robust monitors.
  • Novel Control Protocols: Implementing and testing protocols such as the "U picks from T outputs" approach, untrusted planning with trusted implementation, debate protocols where T and U instances debate code safety, or iterative editing protocols to combat "red herring" attacks.
  • Red Team Strategies: Developing steganography approaches, training better attack policies for producing backdoors, or creating model organisms that demonstrate collusion with monitors.
  • Evaluation Systems: Creating quizzes to test models' knowledge of deceptive strategies, building free-form "interviews" to probe AI capabilities, or developing consistency checks between multiple factorizations of problems.

Support Resources

Mentors from Redwood Research and UK AISI will provide guidance throughout the weekend, helping participants understand the ControlArena framework, design effective experiments, and troubleshoot issues.

Note: most of the suggested Open Problems assume familiarity with the main Control paper

Each team has access to $400 in Lambda Labs computing credits to support their experiments with larger models or more extensive evaluations.

Guidelines

All projects must be submitted by Monday, March 31, 4:00 AM UTC through the official submission form.

Your submission must include:

  1. A completed project report using the provided template
  2. In case, you want to use Overleaf, here is a template for that
  3. Link to a public GitHub repository with your code
  4. A brief (2-3 minute) video demonstration (optional but recommended)

Important: Include an appendix called "Prompts Used" in your submission that documents any LLM prompts used in your project. This helps others understand exactly what was done with the LLM and improves reproducibility.

🏆 Judging Criteria

Projects will be evaluated based on:

1. Innovation & Literature Foundation (30%)

  • Understanding of existing AI control techniques
  • Novel contributions to the field
  • Appropriate use of references and background literature

2. AI Safety Impact (40%)

  • Clear articulation of the relevant threat model
  • Effectiveness of the proposed control mechanism
  • Thoroughness of testing against potential subversion
  • Potential real-world applicability of the approach

3. Technical Quality & Documentation (30%)

  • Implementation quality and technical rigor
  • Reproducibility of results
  • Clarity of documentation
  • Effective visualization and communication of findings

❓ Frequently Asked Questions

About AI Control

Q: What is AI control?
A: AI control refers to techniques that mitigate security risks from AI systems even if the AI itself is trying to subvert them. Unlike alignment research that aims to make AI systems want the right things, control research assumes the system might have problematic objectives and focuses on restricting its capabilities through external mechanisms.

Q: What's the difference between alignment and control?
A: Alignment focuses on ensuring AI systems have the right goals and values. Control addresses the complementary problem: restricting AI capabilities if alignment fails. Control mechanisms are designed to work even against actively deceptive or adversarial AI.

Q: What are some examples of control techniques?
A: Examples include trusted monitoring (using a weaker but more reliable model to check a more capable model's outputs), sandboxing, runtime verification, debate protocols, and bounded optimization approaches.

Q: Is control research only relevant for AGI or superintelligence?
A: No, control techniques are valuable for current AI systems too. For example, they can help prevent data exfiltration, ensure code produced by AI doesn't contain backdoors, and mitigate various security risks from today's models.

About the Hackathon

Q: Do I need prior experience with AI control to participate?
A: No! While familiarity with the concepts helps, we welcome participants from diverse backgrounds. Reading the foundational paper "AI Control: Improving Safety Despite Intentional Subversion" is recommended before the event.

Q: I'm not an ML expert. Can I still contribute?
A: Absolutely! Control research benefits from diverse skills. If you have experience in cybersecurity, formal verification, software engineering, or other adjacent fields, your perspective will be valuable.

Q: What should I bring to the hackathon?
A: Bring your laptop and charger. For in-person participants, a reusable water bottle and any personal items you might need for a weekend of work. Everything else will be provided.

Q: How do teams work? Do I need to form one in advance?
A: Teams can have up to 6 members. You can form teams in advance or join the team matchmaking session on Saturday. Solo participants are also welcome, though collaboration is encouraged.

Q: What resources will be available during the hackathon?
A: Each team will receive $400 in Lambda Labs compute credits. Technical mentors from Redwood Research and UK AISI will be available during office hours. The ControlArena framework provides testbeds for evaluating control techniques.

Q: How technical should our submission be?
A: Submissions should include technical detail sufficient for reproducibility. However, the clarity of communication is also important. Balance technical depth with accessible explanation.

Q: Is there a code of conduct?
A: Yes. All participants must adhere to the hackathon code of conduct, which promotes respectful collaboration, ethical research practices, and responsible AI development.

Q: I can't attend in person. Can I participate virtually?
A: Yes! Virtual participation is welcome and encouraged. Connect through our Discord server to collaborate with teams globally.

Speakers

  • Buck Shlegeris

    Buck Shlegeris

    Speaker and Judge

    Former MIRI researcher turned CEO of Redwood Research, pioneering technical AI safety and control mechanisms while bridging fundamental research with practical applications.

Judges and mentors

Organizers

Local sites

Where a Sprint can lead

How our programs connect
  1. Sprint

    Anyone can join

    Stand out

  2. Apart Fellowship

    6 to 16 weeks on your own project, with a research project manager, compute and publication support.