
Mar 7 - 10, 2025Online and in person
Women in AI Safety Hackathon
Shape the future of safe and ethical AI development! Whether you're a researcher, developer, policy enthusiast, or new to AI safety - join us for an empowering weekend of innovation and collaboration. No prior AI safety experience required. Together, we can build the foundational tools and frameworks needed for responsible AI development.
Entries
- Education track prizeView project: Morph: AI Safety Education Adaptable to (Almost) Anyone
Morph: AI Safety Education Adaptable to (Almost) Anyone
Team Morph
One-liner: Morph is the ultimate operation stack for AI safety education—combining dynamic localization, policy simulations, and ecosystem tools to turn abstract risks into actionable, culturally relevant solutions for learners worldwide. AI safety education struggles with cultural homogeneity, abstract technical …
- Mechanistic Interpretability PrizeView project: Red-teaming with Mech-Interpretability
Red-teaming with Mech-Interpretability
Red teaming large language models (LLMs) is crucial for identifying vulnerabilities before deployment, yet systematically creating effective adversarial prompts remains challenging. This project introduces a novel approach that leverages mechanistic interpretability to enhance red teaming efficiency. We developed a …
- Social Sciences track prizeView project: Detecting Malicious AI Agents Through Simulated Interactions
Detecting Malicious AI Agents Through Simulated Interactions
Team SafeAIGuard
This research investigates malicious AI Assistants’ manipulative traits and whether the behaviours of malicious AI Assistants can be detected when interacting with human-like simulated users in various decision-making contexts. We also examine how interaction depth and ability of planning influence malicious AI …
- View project: AI Through the Human Lens Investigating Cognitive Theories in Machine Psychology
AI Through the Human Lens Investigating Cognitive Theories in Machine Psychology
Team Ak-Rish
We investigate whether Large Language Models (LLMs) exhibit human-like cognitive patterns under four established frameworks from psychology: Thematic Apperception Test (TAT), Framing Bias, Moral Foundations Theory (MFT), and Cognitive Dissonance. We evaluate GPT-4o, QvQ 72B, LLaMA 70B, Mixtral 8x22B, and DeepSeek V3 …
- View project: HalluShield: A Mechanistic Approach to Hallucination Resistant Models
HalluShield: A Mechanistic Approach to Hallucination Resistant Models
Team HalluShield
Our project tackles the critical problem of hallucinations in large language models (LLMs) used in healthcare settings, where inaccurate information can have serious consequences. We developed a proof-of-concept system that classifies LLM-generated responses as either factual or hallucinated. Our approach leverages …
- View project: Attention Pattern Based Information Flow Visualization Tool
Attention Pattern Based Information Flow Visualization Tool
Team babushka’s
Understanding information flow in transformer-based language models is crucial for mechanistic interpretability. We introduce a visualization tool that extracts and represents attention patterns across model components, revealing how tokens influence each other during processing. Our tool automatically identifies and …
- View project: AI Hallucinations in Healthcare: Cross-Cultural and Linguistic Risks of LLMs in Low-Resource Languages
AI Hallucinations in Healthcare: Cross-Cultural and Linguistic Risks of LLMs in Low-Resource Languages
Team Sapentiae
This project explores AI hallucinations in healthcare across cross-cultural and linguistic contexts, focusing on English, French, Arabic, and a low-resource language, Ewe. We analyse how large language models like GPT-4, Claude, and Gemini generate and disseminate inaccurate health information, emphasising the …
- View project: Inspiring People to Go into RL Interp
Inspiring People to Go into RL Interp
Team Interpreters
This project is attempting to complete the Public Education Track, taking inspiration from ideas 1 and 4. The journey mapping was inspired by bluedot impact and aims to create a course that helps explain the need for work to be done in Reinforcement Learning (RL) interp, especially in the problems of reward hacking …
- View project: Debugging Language Models with SAEs
Debugging Language Models with SAEs
Team SAE Mechanic
This report investigates an intriguing failure mode in the Llama-3.1-8B-Instruct model: its inconsistent ability to count letters depending on letter case and grammatical structure. While the model correctly answers "How many Rs are in BERRY?", it struggles with "How many rs are in berry?", suggesting that uppercase …
- View project: Feature-based analysis of cooperation-relevant behaviour in Prisoner’s Dilemma
Feature-based analysis of cooperation-relevant behaviour in Prisoner’s Dilemma
Team PMA
We hypothesise that internal-based model probing and editing might provide higher signal in multi-agent settings. We implement a small simulation of Prisoner’s Dilemma to probe for cooperation-relevant properties. Our experiments demonstrate that feature-based steering highlights deception-relevant features and does …
- View project: Medical Agent Controller
Medical Agent Controller
Team MAC
The Medical Agent Controller (MAC) is a multi-agent governance framework designed to safeguard AI-powered medical chatbots by intercepting unsafe recommendations in real time. It employs a dual-phase approach, using red-team simulations during testing and a controller agent during production to monitor and intervene …
- View project: A Noise Audit of LLM Reasoning in Legal Decisions
A Noise Audit of LLM Reasoning in Legal Decisions
Team Mindy, Markela, May
AI models are increasingly applied in judgement tasks, but we have little understanding of how their reasoning compares to human decision-making. Human decision-making suffers from bias and noise, which causes significant harm in sensitive contexts, such as legal judgment. In this study, we evaluate LLMs on a legal …
- View project: Moral Wiggle Room in AI
Moral Wiggle Room in AI
Team Warwick AI Safety Team
Does AI strategically avoid ethical information by exploiting moral wiggle room?
- View project: Searching for Universality and Equivariance in LLMs using Sparse Autoencoder Found Features
Searching for Universality and Equivariance in LLMs using Sparse Autoencoder Found Features
Team Meru & Jason
The project investigates how neuron features with properties of universality and equivariance affect the controllability and safety of large language models, finding that behaviors supported by redundant features are more resistant to manipulation than those governed by singular features.
- View project: Mechanistic Interpretability Track: Neuronal Pathway Coverage
Mechanistic Interpretability Track: Neuronal Pathway Coverage
Team 42-Shot
Our study explores mechanistic interpretability by analyzing how Llama 3.3 70B classifies political content. We first infer user political alignment (Biden, Trump, or Neutral) based on tweets, descriptions, and locations. Then, we extract the most activated features from Biden- and Trump-aligned datasets, ranking them …
- View project: LLM Military Decision-Making Under Uncertainty: A Simulation Study
LLM Military Decision-Making Under Uncertainty: A Simulation Study
Team Uncertain.ai
LLMs tested in military decision scenarios typically favor diplomacy over conflict, though uncertainty and chain-of-thought reasoning increase aggressive recommendations. This suggests context-specific limitations for LLM-based military decision support.
- View project: AI-Powered Policymaking: Behavioral Nudges and Democratic Accountability
AI-Powered Policymaking: Behavioral Nudges and Democratic Accountability
Team AI watch
This research explores AI-driven policymaking, behavioral nudges, and democratic accountability, focusing on how governments use AI to shape citizen behavior. It highlights key risks such as transparency, cognitive security, and manipulation. Through a comparative analysis of the EU AI Act and Singapore’s AI …
- View project: AI Bias in Resume Screening
AI Bias in Resume Screening
Team Bias_hiring
Our project investigates gender bias in AI-driven resume screening using mechanistic interpretability techniques. By testing a language model's decision-making process on resumes differing only by gendered names, we uncovered a statistically significant bias favoring male-associated names in ambiguous cases. Using …
- View project: An Interpretable Classifier based on Large scale Social Network Analysis
An Interpretable Classifier based on Large scale Social Network Analysis
Team Team99
Mechanistic model interpretability is essential to understand AI decision making, ensuring safety, aligning with human values, improving model reliability and facilitating research. By revealing internal processes, it promotes transparency, mitigates risks, and fosters trust, ultimately leading to more effective and …
- View project: Latent Knowledge Analysis via Feature-Based Causal Tracing
Latent Knowledge Analysis via Feature-Based Causal Tracing
Team Individual submission
This project explores how factual knowledge is stored in large language models using Goodfire’s Ember API. By identifying and manipulating internal features related to specific facts, it shows how facts are encoded and how model behavior changes when those features are amplified or erased.
- View project: Superposition, but at a Cross-MLP Layers view?
Superposition, but at a Cross-MLP Layers view?
Team Snorlax
To understand causal relationships between features (extracted by SAE) across MLP layers, this study introduces the Coordinated Sparse Autoencoder Network (CoSAEN). CoSAEN integrates sparse autoencoders for feature extraction with the PC algorithm for causal discovery, to find the path-based activations of features in …
- View project: Scam Detective: Using Gamification to Improve AI-Powered Scam Awareness
Scam Detective: Using Gamification to Improve AI-Powered Scam Awareness
Team ATBP
This project outlines the development of an interactive web application aimed at involving users in understanding the AI skills for both producing believable scams and identifying deceptive content. The game challenges human players to determine if text messages are genuine or fraudulent against an AI. The project …
- View project: Hikayat - Interactive Stories to Learn AI Safety
Hikayat - Interactive Stories to Learn AI Safety
Team Dune Tech
This paper presents an interactive, scenario-based learning approach to raise public awareness of AI risks and promote responsible AI development. By leveraging Hikayat, traditional Arab storytelling, the project engages non-technical audiences, emphasizing the ethical and societal implications of AI, such as privacy, …
- View project: Preparing for Accelerated AGI Timelines
Preparing for Accelerated AGI Timelines
Team The proactives
This project examines the prospect of near-term AGI from multiple angles—careers, finances, and logistical readiness. Drawing on various discussions from LessWrong, it highlights how entrepreneurs and those who develop AI-complementary skills may thrive under accelerated timelines, while traditional, incremental …
- View project: BUGgy: Supporting AI Safety Education through Gamified Learning
BUGgy: Supporting AI Safety Education through Gamified Learning
Team AI Safety Initiative Groningen
As Artificial Intelligence (AI) development continues to proliferate, educating the wider public on AI Safety and the risks and limitations of AI increasingly gains importance. AI Safety Initiatives are being established across the world with the aim of facilitating discussion-based courses on AI Safety. However, …
- View project: U Reg AI: you regulate it, or you regenerate it!
U Reg AI: you regulate it, or you regenerate it!
Team U Reg AI
We have created a 'choose your path' role game to mitigate existential AI risk ... at this point they might be actual situations in the near-future. The options for mitigation are holistic and dynamic to the player's previous choices. The final result is an evaluation of the player's decision-making performance in …
- View project: Identification if AI generated content
Identification if AI generated content
Team Apollo_team
Our project falls within the Social Sciences track, focusing on the identification of AI-generated text content and its societal impact. A significant portion of online content is now AI-generated, often exhibiting a level of quality and human-likeness that makes it indistinguishable from human-created content. This …
- View project: BlueDot Impact Connect: A Comprehensive AI Safety Community Platform
BlueDot Impact Connect: A Comprehensive AI Safety Community Platform
Team BlueDot Impact Connect
Track: Public Education The AI safety field faces a critical challenge: while formal education resources are growing, personalized guidance and community connections remain scarce, especially for newcomers from diverse backgrounds. We propose BlueDot Impact Connect, a comprehensive AI Safety Community Platform …
- View project: Beyond Statistical Parrots: Unveiling Cognitive Similarities and Exploring AI Psychology through Human-AI Interaction
Beyond Statistical Parrots: Unveiling Cognitive Similarities and Exploring AI Psychology through Human-AI Interaction
Recent critiques labeling large language models as mere "statistical parrots" overlook essential parallels between machine computation and human cognition. This work revisits the notion by contrasting human decision-making—rooted in both rapid, intuitive judgments and deliberate, probabilistic reasoning (System 1 and …
- View project: AI Society Tracker
AI Society Tracker
Team ayoh
My project aimed to develop a platform for real time and democratized data on ai in society
- View project: Interactive Assessments for AI Safety: A Gamified Approach to Evaluation and Personal Journey Mapping
Interactive Assessments for AI Safety: A Gamified Approach to Evaluation and Personal Journey Mapping
Team EtherFlow
An interactive assessment platform and mentor chatbot hosted on Canvas LMS, for testing and guiding learners from BlueDot's Intro to Transformative AI Course.
- View project: SafeAI Academy - Enhancing AI Safety Awareness through Interactive Learning
SafeAI Academy - Enhancing AI Safety Awareness through Interactive Learning
Team SafeAI Solo
SafeAI Academy is an interactive learning platform designed to teach AI safety principles through engaging scenarios and quizzes. By simulating real-world AI challenges, users learn about bias, misinformation, and ethical AI decision-making in an interactive and stress-free environment. The platform uses gamification, …
- View project: AI Safety Escape Room
AI Safety Escape Room
Team NV
The AI Safety Escape Room is an engaging and hands-on AI safety simulation where participants solve real-world AI vulnerabilities through interactive challenges. Instead of learning AI safety through theory, users experience it firsthand – debugging models, detecting adversarial attacks, and refining AI fairness, all …
Overview
Prize Winners
The hackathon demonstrated exceptional innovation across all tracks, with judges particularly impressed by the comprehensive integration of technical rigor with practical AI safety applications. All winning projects showed clear potential for real-world impact and continued development beyond the hackathon weekend.
🥇 Social Sciences Track Winner "Detecting Malicious AI Agents Through Simulated Interactions" by Yulu Pi, Anna Becker, and Ella Bettison
This groundbreaking project introduced Intent-Aware Prompting (IAP) for detecting malicious AI behavior through systematic human-like interaction scenarios. The team developed a comprehensive methodology that revealed how vulnerability increases with interaction depth, providing critical insights for real-world AI deployment contexts and bridging technical research with social science approaches.
🥇 Mechanistic Interpretability Track Winner -"Red-teaming with Mech-Interpretability" by Devina Jain
An innovative fusion of mechanistic interpretability and red-teaming techniques, this project created a real-time dashboard that analyzes neural activation patterns to identify unsafe model behaviors. The system successfully scraped and analyzed jailbreak attempts, identifying "high entropy" prompts most likely to elicit dangerous outputs while enabling targeted safety refinements based on internal model states.
🥇 Public Education Track Winner - $600 "Morph: AI Safety Education Adaptable to (Almost) Anyone" by Shafira Noh and Wan Aimran
This culturally adaptive educational platform addresses the critical gap in global AI safety education by creating personalized learning pathways that bridge theory and practice across diverse cultural contexts. Built on strong literature foundations, Morph tackles the homogeneity problem in AI safety education, making complex concepts accessible to learners worldwide regardless of their technical background or cultural context.
The Women in AI Safety Hackathon brings together talented individuals to tackle crucial challenges in AI development and deployment. This event particularly encourages women and underrepresented groups to contribute their unique perspectives to critical areas of AI safety, including alignment, governance, security, and evaluation.
As AI systems become increasingly powerful and pervasive, diverse perspectives in their development and safety mechanisms are more crucial than ever. This hackathon provides a platform for participants to:
- Collaborate with leading women researchers and practitioners in AI safety
- Develop practical solutions to pressing AI safety challenges
- Build lasting connections in the AI safety community
- Receive mentorship from experienced professionals
- Present ideas to industry experts
Challenge Tracks
- AI Alignment & Values
- Developing methods for value learning
- Improving reward modeling
- Enhancing the interpretability of AI systems
- Safety Evaluations & Testing
- Creating robust testing frameworks
- Developing evaluation metrics
- Building assessment tools for AI systems
- AI Governance & Policy
- Designing accountability frameworks
- Developing oversight mechanisms
- Creating safety standards
- Technical AI Safety
- Addressing mesa-optimization
- Working on robustness
- Developing safety architectures
Resources
Schedule
The schedule runs from 4 PM UTC Friday to 3 AM Monday. We start with an introductory talk and end the event during the following week with an awards ceremony. Join the public ICal here. You will also find Explorer events, such as collaborative brainstorming and team match-making, before the hackathon begins on Discord and in the calendar.
Speakers
Judges and mentors
Organizers
- (opens in new tab)

Natalia Pérez-Campanero Antolín
Organiser

Myra Deng
Organizer
- (opens in new tab)

Lindsey Robertson
Co-Organizer
- (opens in new tab)

Grecia Castaldi
Co-Organizer
- (opens in new tab)

ChengCheng Tan
Co-Organizer
- (opens in new tab)

Ziba Atak
Co-Organizer
- (opens in new tab)

Archana Vaidheeswaran
Organizer
- (opens in new tab)

Angela Cao
Co-Organizer

Andreea Damien
Social Science Track Organiser and Judge
- (opens in new tab)

Astha Puri
Organiser
Local sites
AI Safety Hackathon (by WAI and EA Warwick)
Join Warwick AI and Effective Altruism to the joint hackathon Weekend. We will be on the 8th and 9th March in room FAB 2.48 on the main Campus of the Universtiy of Warwick. Anyone can join! Hope to see you there!
Event page: AI Safety Hackathon (by WAI and EA Warwick) (opens in new tab)AISIG - Women in AI Safety Research Hackathon
Join us for the Women in AI Safety Research Hackathon in Hereplein 4, 9711GA, Groningen!
Event page: AISIG - Women in AI Safety Research Hackathon (opens in new tab)Women in AI Safety Hackathon
25 Holywell Row, London EC2A 4XE
Event page: Women in AI Safety Hackathon (opens in new tab)Women in AI Safety Hackathon - 42AI PARIS jam site
42AI is a student association dedicated to foster learning and discussion in the field of AI.
Event page: Women in AI Safety Hackathon - 42AI PARIS jam site (opens in new tab)Women in AI Safety Hackathon - Dubai
If you're coming by Taxi or Metro, Enter the Boulevard area of Emirates Towers. You'll find CodersHQ right opposite Creators Hub on the Ground Floor. If you're driving, there is free parking here: https://maps.app.goo.gl/pp1nqVbrN2LiQbk48
Event page: Women in AI Safety Hackathon - Dubai (opens in new tab)Women in AI Safety Research Hackathon
Join us at the EA Hotel for the Women in AI Safety Research Hackathon. Free accommodation, food and co-working stations provided! We're located at: 36 York Street, Blackpool, FY15AQ. Please register through our Luma event page so we know you're coming!
Event page: Women in AI Safety Research Hackathon (opens in new tab)
Where a Sprint can lead
How our programs connectAnyone can join
Stand out
6 to 16 weeks on your own project, with a research project manager, compute and publication support.
Upcoming Sprints
All SprintsAI Collusion Research Sprint
A weekend research sprint on collusion between AI agents: when it emerges in markets and everyday workflows, how to detect and audit it, how it is carried, and what breaks it. Co-organized with Poseidon Research and AE Studio, online with in-person hubs at Collider in New York City and AI Safety Hong Kong. Top teams are invited to apply to the Apart Fellowship.
Read the brief: AI Collusion Research SprintAI x Epistemics Research Sprint
A weekend research sprint on AI for epistemics: evaluating whether models know how solid their claims are, building trust infrastructure that people and agents can consume, and shipping epistemic products that improve real decisions. Online, four tracks including an open track. Top teams are invited to apply to the Apart Fellowship.
Read the brief: AI x Epistemics Research SprintQuestions? sprints@apartresearch.com







