Skip to content
Women in AI Safety Hackathon

Mar 7 - 10, 2025Online and in person

Women in AI Safety Hackathon

This Sprint has ended.

Sign-ups
334
Projects submitted
33
Browse the 33 projects

Sign up for this Sprint

Type N/A if you don’t have one.

Type N/A if you don’t have one.

What you work on, and whether you are open to new roles.

What about this event made you want to take part?

By signing up you agree to our Privacy Policy.

The submission window has closed. Your draft is still here so you can copy it, but it can no longer be submitted.

Submit your project

Project details

A short abstract: what you did, what you found.

PDF, up to 25 MB.

Are you interested in publishing this project? *
Tracks

If this Sprint has numbered tracks, choose the ones your project fits.

PDF, PowerPoint, Keynote or ODP, up to 25 MB.

PNG, JPEG, WebP or GIF, up to 25 MB.

Team details

Team member 1

Leave blank if you don’t have one.

By submitting you agree to the prize terms and our Privacy Policy.

The submission window has closed. Your draft is still here so you can copy it, but it can no longer be submitted.

See upcoming Sprints

Shape the future of safe and ethical AI development! Whether you're a researcher, developer, policy enthusiast, or new to AI safety - join us for an empowering weekend of innovation and collaboration. No prior AI safety experience required. Together, we can build the foundational tools and frameworks needed for responsible AI development.

Entries

Overview

Prize Winners

The hackathon demonstrated exceptional innovation across all tracks, with judges particularly impressed by the comprehensive integration of technical rigor with practical AI safety applications. All winning projects showed clear potential for real-world impact and continued development beyond the hackathon weekend.

🥇 Social Sciences Track Winner "Detecting Malicious AI Agents Through Simulated Interactions" by Yulu Pi, Anna Becker, and Ella Bettison

This groundbreaking project introduced Intent-Aware Prompting (IAP) for detecting malicious AI behavior through systematic human-like interaction scenarios. The team developed a comprehensive methodology that revealed how vulnerability increases with interaction depth, providing critical insights for real-world AI deployment contexts and bridging technical research with social science approaches.

🥇 Mechanistic Interpretability Track Winner -"Red-teaming with Mech-Interpretability" by Devina Jain

An innovative fusion of mechanistic interpretability and red-teaming techniques, this project created a real-time dashboard that analyzes neural activation patterns to identify unsafe model behaviors. The system successfully scraped and analyzed jailbreak attempts, identifying "high entropy" prompts most likely to elicit dangerous outputs while enabling targeted safety refinements based on internal model states.

🥇 Public Education Track Winner - $600 "Morph: AI Safety Education Adaptable to (Almost) Anyone" by Shafira Noh and Wan Aimran

This culturally adaptive educational platform addresses the critical gap in global AI safety education by creating personalized learning pathways that bridge theory and practice across diverse cultural contexts. Built on strong literature foundations, Morph tackles the homogeneity problem in AI safety education, making complex concepts accessible to learners worldwide regardless of their technical background or cultural context.

The Women in AI Safety Hackathon brings together talented individuals to tackle crucial challenges in AI development and deployment. This event particularly encourages women and underrepresented groups to contribute their unique perspectives to critical areas of AI safety, including alignment, governance, security, and evaluation.

As AI systems become increasingly powerful and pervasive, diverse perspectives in their development and safety mechanisms are more crucial than ever. This hackathon provides a platform for participants to:

  • Collaborate with leading women researchers and practitioners in AI safety
  • Develop practical solutions to pressing AI safety challenges
  • Build lasting connections in the AI safety community
  • Receive mentorship from experienced professionals
  • Present ideas to industry experts

Challenge Tracks

  1. AI Alignment & Values
    • Developing methods for value learning
    • Improving reward modeling
    • Enhancing the interpretability of AI systems
  2. Safety Evaluations & Testing
    • Creating robust testing frameworks
    • Developing evaluation metrics
    • Building assessment tools for AI systems
  3. AI Governance & Policy
    • Designing accountability frameworks
    • Developing oversight mechanisms
    • Creating safety standards
  4. Technical AI Safety
    • Addressing mesa-optimization
    • Working on robustness
    • Developing safety architectures

Schedule

The schedule runs from 4 PM UTC Friday to 3 AM Monday. We start with an introductory talk and end the event during the following week with an awards ceremony. Join the public ICal here. You will also find Explorer events, such as collaborative brainstorming and team match-making, before the hackathon begins on Discord and in the calendar.

Speakers

  • Zainab Majid

    Zainab Majid

    Speaker

    Zainab works at the intersection of AI safety and cybersecurity, leveraging her expertise in incident response investigations to tackle AI security challenges.

  • Tarin Rickett

    Tarin Rickett

    Speaker

    Product & Engineering Lead at BlueDot Impact, former LinkedIn Staff Engineer. CS and Brain Science grad from Rochester, passionate about educational tech and women in computing.

Judges and mentors

Organizers

Local sites

Where a Sprint can lead

How our programs connect
  1. Sprint

    Anyone can join

    Stand out

  2. Apart Fellowship

    6 to 16 weeks on your own project, with a research project manager, compute and publication support.