Skip to content
The Technical AI Governance Challenge

Jan 30 - Feb 1, 2026Online

The Technical AI Governance Challenge

This Sprint has ended.

Sign-ups
341
Projects submitted
47
Browse the 47 projects

Sign up for this Sprint

Type N/A if you don’t have one.

Type N/A if you don’t have one.

What you work on, and whether you are open to new roles.

What about this event made you want to take part?

By signing up you agree to our Privacy Policy.

The submission window has closed. Your draft is still here so you can copy it, but it can no longer be submitted.

Submit your project

Project details

A short abstract: what you did, what you found.

PDF, up to 25 MB.

Are you interested in publishing this project? *
Tracks

Choose every track your project fits.

PDF, PowerPoint, Keynote or ODP, up to 25 MB.

PNG, JPEG, WebP or GIF, up to 25 MB.

Team details

Team member 1

Leave blank if you don’t have one.

By submitting you agree to the prize terms and our Privacy Policy.

The submission window has closed. Your draft is still here so you can copy it, but it can no longer be submitted.

See upcoming Sprints

This hackathon brings together 500+ builders to prototype verification tools, compliance systems, and coordination infrastructure that enable international cooperation on frontier AI safety.

Entries

Overview

HACKATHON WINNERS

Huge congratulations to all our winners, and thank you to everyone who participated and to our incredible panel of 17 judges who reviewed every single project. With 300+ participants and 48 projects submitted, the competition was fierce and the projects were outstanding. Here's who came out on top:

—————————————————————————————————————————————————

Frontier AI labs are training systems that could pose international risks. Countries want agreements on safe development. Labs need ways to demonstrate compliance without exposing competitive advantages.

The technical infrastructure to make this possible doesn't exist yet. We have policy frameworks without verification systems. International agreements without monitoring tools. Compliance requirements without practical implementation paths.

This hackathon focuses on building that missing infrastructure. You'll have one intensive weekend to create verification protocols, monitoring tools, privacy-preserving compliance proofs, or coordination systems that could enable enforceable international cooperation on AI safety.

Top teams get:

💰 $2000 in cash prizes + Fast track to:
Your Next Job at Lucid Computing
Product Engineer
Hardware Security
Compiler & Kernel
Network Security
The MIRI TGT Fellowship
The Apart Fellowship

Fast-tracks include at least one interview with leadership from Lucid Computing, or researchers at MIRI TGT or Apart Research.

What is International Technical AI Governance?

Technical international governance refers to the practical infrastructure needed to verify, monitor, and enforce international agreements on AI development. This includes:

  • Hardware verification that tracks compute resources used in training frontier models
  • Attestation systems that cryptographically prove model properties without revealing weights
  • Privacy-preserving compliance proofs using zero-knowledge cryptography or trusted execution environments
  • Risk threshold frameworks that define when capabilities trigger safety requirements
  • International coordination mechanisms that enable verification between parties without full trust
  • Dual-use detection for identifying dangerous capabilities in AI research before deployment

Labs are training increasingly capable systems. Some pose risks that cross borders. International cooperation requires technical mechanisms to verify compliance without exposing sensitive information or creating security vulnerabilities.

Why this hackathon?

The Problem

AI systems get more capable, our governance infrastructure doesn't. Multiple labs are training models above the EU's 10²⁵ FLOP threshold for systemic risk. Some frontier models now require ASL-3 safeguards. Decentralized training makes compute monitoring harder to implement.

The EU AI Act took effect in August 2024, but practical compliance tools remain scarce. Export controls on AI chips lack verification mechanisms. Responsible scaling policies define thresholds without automated monitoring. Labs coordinate through voluntary frameworks that lack enforcement infrastructure.

Most governance proposals assume technical capabilities that don't exist yet. They require compute tracking systems not deployed at scale, attestation mechanisms not integrated into hardware, and verification protocols not tested between adversarial parties. We're building policy without infrastructure.

Why International Technical AI Governance Matters Now

International cooperation depends on verifiable compliance. If agreements can't be verified without exposing sensitive IP or creating security risks, countries won't sign them. Labs won't share information that compromises their competitive position.

We're massively under-investing in governance infrastructure. Most effort goes into capabilities research or post-deployment harm mitigation. Far less into building the verification systems, monitoring tools, and coordination mechanisms that enable international agreements.

Better technical infrastructure could give us agreements that labs can verify without exposing model weights, monitoring systems that respect privacy while enabling compliance checks, and coordination mechanisms that work between parties without full trust. It could create the practical foundation needed for international cooperation on frontier AI safety.

Hackathon Tracks

1. Hardware Verification & Attestation

  • Design hardware verification protocols for tracking compute resources in datacenter environments
  • Build attestation systems using trusted execution environments (TEEs) that prove model properties without exposing weights
  • Create compute monitoring tools that detect training runs above regulatory thresholds
  • Develop chip-level security mechanisms for remote verification of AI hardware properties

2. Compliance Infrastructure & Privacy-Preserving Proofs

  • Build zero-knowledge proof systems that demonstrate regulatory compliance without revealing sensitive information
  • Create privacy-preserving audit mechanisms for federated learning or distributed training
  • Develop compliance automation tools for EU AI Act requirements, GPAI reporting, or safety frameworks
  • Design cryptographic protocols that enable verification between parties without full trust

3. Risk Thresholds & Compute Verification

  • Build risk assessment frameworks that map compute thresholds to capability levels
  • Create tools for harmonizing ASL/CCL terminology across different lab safety frameworks
  • Develop capability evaluation systems for dual-use risks (CBRN, cyber, autonomous AI R&D)
  • Design monitoring systems for responsible scaling policies and deployment safeguards

4. International Verification & Coordination

  • Build coordination infrastructure for International Network of AI Safety Institutes
  • Create verification mechanisms inspired by IAEA frameworks adapted for AI governance
  • Develop systems for cross-border information sharing that respect national security concerns
  • Design tools for implementing global AI safety standards and red lines

5. Research Governance & Dual-Use Detection

  • Build detection systems for identifying dangerous capabilities in pre-publication research
  • Create frameworks for assessing dual-use risks in biological AI models or other specialized domains
  • Develop pre-publication review tools that scale across research communities
  • Design capability-based threat assessment systems for frontier AI research

Who should participate?

This hackathon is for people who want to build solutions to technological risk using technology itself.

You should participate if you're an engineer, researcher, or developer who wants to work on consequential problems and build practical verification, monitoring, or compliance infrastructure.

No prior governance research experience required. We provide resources, mentors, and starter templates.

What you will do

Participants will:

  • Form teams or join existing groups.
  • Develop projects over an intensive hackathon weekend.
  • Submit open-source verification tools, compliance systems, monitoring infrastructure, or empirical research advancing international AI governance

Please note: Due to the high volume of submissions, we cannot guarantee written feedback for every participant, although all projects will be evaluated.

What happens next

Winning and promising projects will be:

  • Awarded $2,000 in cash prizes
  • Fast-tracked for interviews with Lucid Computing, MIRI TGT, or Apart Research
  • Published openly for the community
  • Invited to continue development within the Apart Fellowship
  • Shared with relevant safety researchers and policymakers.

Why join?

  • Work on consequential problems: Build infrastructure that could enable international cooperation on frontier AI safety
  • Learn from experts: Get guidance from AI safety researchers and technical governance practitioners throughout the weekend
  • Build your network: Collaborate with technical talent from across the globe focused on AI safety
  • Develop practical skills: Gain hands-on experience with verification systems, cryptographic proofs, or monitoring infrastructure that employers value

Resources

General Introduction

  • International AI Safety Report 2025
    Led by Yoshua Bengio, backed by 30 countries and international organizations
    The inaugural comprehensive scientific review of general-purpose AI capabilities and risks. Essential reading that establishes the evidence base for AI governance discussions, covering capability assessments, risk taxonomies, and technical approaches to safety. Participants will gain a shared vocabulary and understanding of the threat landscape that underpins all hackathon tracks.
  • Computing Power and the Governance of AI
    Centre for the Governance of AI (GovAI)
    The foundational paper explaining why compute is uniquely governable compared to other AI inputs (data, algorithms). Covers compute's detectability, excludability, quantifiability, and supply chain concentration. Essential for understanding why hardware-focused governance is feasible and how visibility, allocation, and enforcement mechanisms work in practice.
  • The Annual AI Governance Report 2025: Steering the Future of AI
    International Telecommunication Union (ITU)
    Comprehensive overview of global AI governance approaches, from Europe's risk-based AI Act to Asia's innovation-driven models. Covers the transition from principles to operational tools, regional variations in governance philosophy, and the emerging role of international coordination. Provides crucial context on the political landscape participants will be building for.

Track 1: Hardware Verification, Attestation & Lifecycle Security

  • Technology to Secure the AI Chip Supply Chain: A Working Paper
    Center for a New American Security (CNAS) – April 2025
    Detailed primer on Hardware-Enabled Mechanisms (HEMs) including location verification, offline licensing, and workload attestation. Explains how these mechanisms could enable targeted export controls, privacy-preserving compliance reporting, and enforcement of international agreements. Essential reading for understanding the technical building blocks of chip governance.
  • Hardware-Enabled Mechanisms for Verifying Responsible AI Development
    arXiv – April 2025
    Technical deep-dive into location verification, offline licensing, workload classification, and detailed verification approaches. Covers open challenges including anti-tamper techniques, privacy protections, and cluster configuration flexibility. Includes practical discussion of Trusted Execution Environments (TEEs) and remote attestation.
  • Flexible Hardware-Enabled Guarantees (FlexHEG) Report
    Future of Life Institute – January 2025
    Ambitious proposal for a family of hardware mechanisms consisting of secure processors within tamper-resistant enclosures that locally enforce flexible policies. Addresses privacy-preserving verification, mutual verification between geopolitical rivals, and the potential for international agreements. Forward-looking vision of what mature chip governance could look like.

Supplementary Resources

Track 2: Compliance Infrastructure, Monitoring & Privacy-Preserving Proofs

  • How the EU's Code of Practice Advances AI Safety
    AI Frontiers – July 2025
    Explains the EU Code of Practice's requirements including risk estimation, external evaluation, incident reporting, and public transparency. Critical for understanding what compliance actually requires under the first major frontier AI regulation—and therefore what compliance infrastructure must support.
  • AI Lab Watch – Commitments Tracker
    Zach Stein-Perlman
    Comprehensive tracking of AI company commitments, from the Seoul Summit pledges to responsible scaling policies. Documents the gap between stated commitments and actual implementation. Essential context on what independent monitoring looks like today and where gaps exist (this effort is winding down, creating an important gap).
  • AI Safety Index Winter 2025
    Future of Life Institute – December 2025
    Systematic evaluation of 7 leading AI companies across 33 indicators of responsible AI development. Provides methodology for assessing compliance with safety commitments, covering risk management, governance, transparency, and disclosure. Model for what rigorous independent monitoring looks like.

Supplementary Resources:

Track 3: Risk Thresholds, Modeling & Compute Verification

  • Common Elements of Frontier AI Safety Policies
    METR (Model Evaluation & Threat Research)
    Comprehensive analysis of capability thresholds, risk tiers, and safeguards across 12 published frontier safety policies from OpenAI, Anthropic, DeepMind, xAI, Amazon, and others. Essential for understanding how labs currently define "dangerous" and where approaches differ. Includes direct quotes from each framework.
  • AI Safety under the EU AI Code of Practice
    Georgetown CSET – July 2025
    Analysis of how the Code of Practice sets a "minimum standard for appropriate risk management" that goes beyond current industry practices. Covers pre-defined risk tiers, capability-based thresholds, and the requirement for external evaluation. Critical for understanding what risk threshold harmonization might look like.
  • The Role of Compute Thresholds for AI Governance
    Institute for Law & AI – February 2025
    Deep analysis of how compute thresholds function as regulatory triggers, their limitations (algorithmic progress, post-training enhancements), and verification challenges. Discusses how "on-chip governance mechanisms" could verify compute claims. Essential for understanding threshold-based governance.

Supplementary Resources:

Track 4: International Verification & Coordination Infrastructure

  • Do We Want an "IAEA for AI"?
    Lawfare – November 2024
    Careful analysis of whether and how IAEA models apply to AI governance. Covers IAEA's monitoring and verification functions, its limitations (North Korea, Iran), and unique challenges for AI (software copyability, compute repurposing, ephemeral training runs). Essential framing for international verification discussions.
  • The Global Landscape of AI Safety Institutes
    All Tech Is Human – May 2025
    Comprehensive catalogue of national AI Safety Institutes worldwide, their functions, and the International Network of AI Safety Institutes. Analyzes the tension between national sovereignty and international coordination, and the recent UK rebranding to "AI Security Institute." Essential map of the institutional landscape.
  • Mechanisms to Verify International Agreements About AI Development
    ResearchGate – June 2025
    Technical paper on low-tech and high-tech approaches to international verification. Covers self-reporting with inspection, on-chip mechanisms for remote attestation, and workload classification for detecting large training runs. Practical roadmap for what near-term verification could look like.

Supplementary Resources:

Track 5: Research Governance & Dual-Use Detection

  • Dual-use Capabilities of Concern of Biological AI Models
    PLOS Computational Biology – May 2025
    Technical analysis of biosecurity risks from AI, including prevention and mitigation strategies: data exclusion, machine unlearning, API access restrictions, and governmental risk-benefit assessment. Model for how dual-use research governance discussions apply to AI-specific contexts.
  • Framework for Artificial Intelligence Diffusion
    Bureau of Industry and Security – January 2025
    The Biden administration's framework for controlling AI chip exports and model weights. Establishes the first regulatory framework treating AI model weights as controlled items alongside hardware. Essential context for understanding export control approaches to dual-use AI.
  • Regulating Artificial Intelligence: U.S. and International Approaches
    Congressional Research Service – 2025
    Comprehensive overview of U.S. AI regulatory approaches including compute thresholds, dual-use reporting requirements, and the policy shift under the Trump administration. Provides essential context on the regulatory landscape and open questions for Congress.

Supplementary Resources

Useful Datasets, Benchmarks & Tools

Datasets

Benchmarks & Evaluations

Open-Source Tools

  • Three.js – 3D visualization (for hardware/network simulations)
  • D3.js – Data visualization
  • LangChain/LlamaIndex – For AI-powered document analysis
  • Plotly – Interactive charts

Project Ideas

Project Scoping Advice

Based on successful hackathon retrospectives:

  1. Focus on MVP, Not Production. In 2 days, aim for:
    1. Day 1: Set up environment, implement core functionality, get basic pipeline working
    2. Day 2: Add 1-2 key features, create demo, prepare presentation
  2. Use Mock/Simulated Data rather than integrating real APIs or databases, use:
    1. Synthetic chip registries or compliance records
    2. Simulated protocol interactions (e.g., attestation handshakes)
    3. Pre-compiled policy documents and model cards
      This eliminates authentication, rate limiting, and data quality issues.
  3. Leverage Existing Frameworks. Don't build from scratch. Use:
    1. Published policy texts (EU Code of Practice, lab RSPs) as structured inputs
    2. Epoch AI datasets for compute trends
    3. AI Incident Database for documented cases
    4. Existing comparison frameworks as starting points
  4. Clear Success Criteria. Define what "working" means:
    1. For hardware verification: Models 3+ attack scenarios with documented threat assumptions
    2. For compliance tools: Tracks 5+ lab commitments with structured change detection
    3. For threshold analysis: Compares definitions across 4+ major labs with gap taxonomy
    4. For international verification: Maps 10+ arms control precedents to AI-specific challenges
    5. For research governance: Classifies 50+ papers with documented methodology and error cases

Guidelines

🏆 Judging Criteria

Dimension 1: Impact Potential & Innovation

How much would this matter for AI safety if it worked? How innovative is it?
For scores of 4-5: is this actually new to the field, or replicating recent work?

ScoreDescription
1Negligible. No clear problem addressed, or no meaningful novelty.
2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

Dimension 2: Execution Quality

How sound are methodology, implementation, and findings?

ScoreDescription
1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

Dimension 3: Presentation & Clarity

How clearly are work, findings, and impact potential communicated?

ScoreDescription
1Incomprehensible. Cannot determine what the project is actually claiming or doing.
2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

Submission Requirements

All projects must be submitted by the deadline through the official submission portal.

Your submission must include:

  • A completed project report using the provided template (mandatory)
  • Link to a public GitHub repository with your code (recommended)
  • A brief (3-5 minute) video demonstration of your solution (optional)
  • An appendix documenting any AI/LLM prompts used in your project for reproducibility (optional)

Important: Include an appendix called "Limitations & Dual-Use Considerations" that addresses the following:

  • Limitations (false positives/negatives, edge cases, scalability constraints)
  • Dual-use risks (could your method be used to train better manipulators?)
  • Responsible disclosure recommendations (if vulnerabilities discovered)
  • Ethical considerations in your approach
  • Suggestions for future improvements

Please note: Due to the high volume of submissions, we cannot guarantee written feedback for every participant, although all projects will be evaluated.

❓ Frequently Asked Questions

About the Technical International Governance Hackathon

Q: Who can participate?
A: Anyone with a strong interest in AI safety research, technical solutions, or AI policy. This includes researchers, engineers, students, policy analysts, and domain experts in other domains.

Q: Do I need AI safety or research experience?
A: No. We provide starter resources and mentors.

Q: Is the hackathon remote?
A: Yes. Global and virtual. Talks are streamed. Collaboration happens on the Apart community Discord

Q: How do teams work?
A: Teams can have up to 5 members. You can form teams in advance or join the team-matching session at the beginning of the event. Solo participants are welcome, though collaboration is encouraged.

Q: What computing resources will be available?
A: Each team will receive $400 in cloud computing credits.*

Q: Is there a code of conduct?
A: Yes. All participants must adhere to the hackathon code of conduct, which promotes responsible research, ethical AI development, and respectful collaboration.

*Cloud compute access to A100s or stronger GPUs is not available to participants from countries with active U.S. sanctions. A list of sanctioned countries can be found here.

Schedule

🗓️
🗓️

FULL SCHEDULE

Wednesday, January 28
10:00 PT - HackTalk: Charbel-Raphaël Segerie, Executive Director at CeSIA
Thursday, January 29
10:00 PT - HackTalk: Henry Papadatos, Executive Director at SaferAI
Friday, January 30
10:00 PT - Keynote: Kristian Rönn, Cofounder of Lucid Computing
18:00 PT - HackTalk (In-Person): Peter Barnett, MIRI Technical Governance Team
19:00 PT - Hacking Begins
Sunday, February 1
23:59 PT - Submission Deadline!

Note: All times listed in PT (Pacific Time, UTC-8)

Speakers

  • Kristian Rönn

    Kristian Rönn

    Keynote Speaker

    Kristian Rönn is the CEO and co-founder of Lucid Computing, an AI hardware governance company building verification infrastructure for compute export controls. Before pivoting to AI safety, he spent 11 years building Normative, a carbon accounting platform that became a leading tool for corporate emissions tracking. His path to tech entrepreneurship started at Oxford's Future of Humanity Institute, where he worked on global catastrophic risks. He's the author of The Darwinian Trap, which examines how evolutionary pressures shape systemic risks to humanity's future.

  • Charbel-Raphaël Segerie

    Charbel-Raphaël Segerie

    Speaker

    Charbel-Raphaël Segerie is the Executive Director of CeSIA, France's leading AI safety research organization. He created Europe's first AI safety course for general-purpose models at ENS Paris-Saclay and founded ML4Good, a bootcamp series that has trained researchers across six countries. His technical work focuses on RLHF limitations and interpretability. He led the Global Call for AI Red Lines, an international campaign signed by 10 Nobel laureates and introduced at the UN General Assembly, and contributes to the EU AI Office's Code of Practice for general-purpose AI systems.

  • Henry Papadatos

    Henry Papadatos

    Speaker

    Henry Papadatos is the Executive Director of SaferAI, where he works on technical solutions for frontier AI risk management. He contributed to the EU AI Act's Codes of Practice as part of the expert working group on risk taxonomy and assessment, and helped draft the G7 Hiroshima AI Process reporting framework through the OECD task force. His technical work includes an AI risk management ratings system for developers and current research on quantitative risk modeling for AI-enabled cyber threats. Before SaferAI, he conducted alignment research on large language models at UC Berkeley's Center for Human-Compatible AI.

  • Peter Barnett

    Peter Barnett

    Speaker

    Peter Barnett is a Technical AI Governance Researcher at the Machine Intelligence Research Institute, where he focuses on preventing catastrophic and extinction risks from artificial intelligence. He co-authored MIRI's international agreement proposal to prevent premature development of artificial superintelligence, which includes verification mechanisms for AI chip usage and training restrictions. His work includes developing AI governance research agendas and addressing trust challenges in multilateral AI agreements. Before MIRI, he conducted alignment research at UC Berkeley's Center for Human-Compatible AI and holds a Master's degree in Physics from the University of Otago, where he specialized in quantum optics simulations.

Organizers

Local sites

Where a Sprint can lead

How our programs connect
  1. Sprint

    Anyone can join

    Stand out

  2. Apart Fellowship

    6 to 16 weeks on your own project, with a research project manager, compute and publication support.