
Nov 21 - 23, 2025Online
Defensive Acceleration Hackathon
This hackathon brings together builders to prototype defensive systems that could protect us from AI-enabled threats.
Entries
- View project: Neops - DevSecOps for the AI era
Neops - DevSecOps for the AI era
Team Broad Bros · London
NEOps is a CLI-based tool that embeds AI safety into your product lifecycle from day one. While development teams routinely build cybersecurity checks, AI-safety often comes later—or not at all. NEOps fills that gap by automatically scanning your code, front-end/back-end architecture, and documentation for threats …
- View project: Assisted Audit of Solana Programs
Assisted Audit of Solana Programs
Team GLAM · San Francisco
Multi-agent solution that assists in auditing Solana programs, allows to consolidate audit findings into a knowledge base, and can integrate into CI/CD pipelines to prevent security regressions.
- View project: Mechanistic Watchdog
Mechanistic Watchdog
Team SL5 · México
Mechanistic Watchdog is a mechanistic-interpretability-based “cognitive kill switch” for language models. Instead of only filtering final text, we monitor a model’s internal activations in real time and learn linear “concept vectors” that capture truthfulness and high-risk domains. Using datasets like …
- View project: GUARDIAN: Guarded Universal Architecture for Defensive Interpretation And traNslation
GUARDIAN: Guarded Universal Architecture for Defensive Interpretation And traNslation
Team Guardian team
GUARDIAN is a multi-stage, LLM-driven system to automate the translation of C codebases to memory-safe Rust. GUARDIAN promotes defense acceleration at-scale by guiding an LLM transpiler with dependency graph strongly-connected-components, static-analysis-guided rule hints, examples from the demonstration corpora and …
- View project: Comparative LLM methods for Social Media Bot Detection
Comparative LLM methods for Social Media Bot Detection
Team Bot Detection · Remote (California, Denmark)
This project examines the potential of LLMs to detect social media bots in (near) real-time, and the potential of using less-advanced LLMs to detect more advanced LLMs. It contributes to the cognitive defense toolset for protecting information ecosystems.
- View project: Durinn Calibration
Durinn Calibration
Team Durinn · Oslo
Durinn analyzes Hacktoberfest repositories to build OWASP-aligned datasets and calibrate a ProtectAI prompt-injection model for detecting vulnerable code. Calibration improves the model’s ability to identify unsafe snippets.
- View project: Robust LLM Neural Activation-Mediated Alignment
Robust LLM Neural Activation-Mediated Alignment
Team Rice AI Alignment · Houston
Large language models (LLMs) can inadvertently generate harmful biological, chemical, or cyber-physical guidance, yet current safety systems rely almost entirely on surface-level text defenses, such as keyword filters, refusal heuristics, or post-generation classifiers, that remain brittle under paraphrasing or …
- View project: Actions speak louder than words: Evaluating Tool Usage Risk in Open-Weight AI for Defensive Deployment
Actions speak louder than words: Evaluating Tool Usage Risk in Open-Weight AI for Defensive Deployment
Team Actions speak louder than words · Spain
Open-weight language models are increasingly deployed in agentic workflows where they can invoke external tools, creating new attack vectors beyond traditional text generation. We present a systematic evaluation framework that measures how model tampering affects tool-usage behavior across cybersecurity, …
- View project: A Defensive AI Agent Against Large Language Model (LLM)-Assisted Polymorphic Malware
A Defensive AI Agent Against Large Language Model (LLM)-Assisted Polymorphic Malware
Team BlueFlux
The rapid evolution of Large Language Models (LLMs) has introduced a new asymmetric threat: AI-assisted polymorphic malware. As identified by Google’s Threat Intelligence Group, attackers are utilizing automated frameworks like "PromptFlux" to weaponize LLMs, generating hundreds of functional, unique malware variants …
- View project: Opening Doors to Multimodal Deception
Opening Doors to Multimodal Deception
Team Happy VLM · Cambridge
Organizations increasingly deploy vision-language models (VLMs) locally to protect sensitive data, yet there is almost no empirical work on whether these multimodal systems can systematically lie about visual content or whether text-based safety tools transfer to visual contexts. We present, to our knowledge, the …
- View project: Automating Privacy-Preserving Model Deployment
Automating Privacy-Preserving Model Deployment
Team 1 · Sydney
- View project: TheWizard
TheWizard
Team Oz Labs · Tel Aviv
TheWizard is a tiny cyber world model built to defend systems against AI agents. TheWizard constructs a symbolic digital twin of an environment and simulates how agents plan and act—files touched, logs queried, credentials accessed—to generate 10–20 plausible futures from any partial trajectory. If an agent’s …
- View project: LLM-ExecGuard - Real-Time Detection of Malicious Shell Behavior in LLM Agents
LLM-ExecGuard - Real-Time Detection of Malicious Shell Behavior in LLM Agents
Team Heron friends · Tel aviv
LLM-ExecGuard is a real-time monitoring system designed to detect malicious behavior by LLM-based agents operating in shell environments. It observes the commands an agent executes and evaluates them using Sigma rules mapped to MITRE ATT&CK tactics. Our prototype hooks into the terminal stream, logs commands outside …
- View project: Detecting Piecewise Cyber Espionage in Model APIs
Detecting Piecewise Cyber Espionage in Model APIs
Team Detecting Piecewise Cyber Espionage in Model APIs · Washington D.C.
Detecting Piecewise Cyber Espionage in Model APIs - On November 13th 2025, Anthropic published a report on an AI-orchestrated cyber espionage campaign. Threat actors used various tech- niques to circumvent model safeguards and used Claude Code with agentic scaffolding to automate large parts of their campaigns. …
- View project: MODX - Inference Time Detection of Anomalous Behavior using Sparse Auto-Encoders
MODX - Inference Time Detection of Anomalous Behavior using Sparse Auto-Encoders
Team modx · Redwood City, CA
The fast adoption of open-source language models in building agents and workflows has amplified the risks of backdoor attacks. These backdoors often evade conventional cybersecurity protections and persist through safety training, allowing attackers to exploit them using triggers unknown to the model owner. We present …
- View project: DuneBox - Prompt Injection Detection with SLM in Local Sandbox
DuneBox - Prompt Injection Detection with SLM in Local Sandbox
Team DUNE · Toronto, Canada
A locked-down sandbox environment designed to safely evaluate language models (SLMs/LLMs) and identify malicious attacks within a Chrome extension. The sandbox isolates model interactions from the broader browser and system, enabling users to test potentially malicious prompts without risking data exposure or …
- View project: From Hallucinations to Misalignment: Evaluating EDFL as a Misalignment Checker on GPT-4o-mini and Sleeper Agents
From Hallucinations to Misalignment: Evaluating EDFL as a Misalignment Checker on GPT-4o-mini and Sleeper Agents
Team The Misaligners · Nottingham, UK
Large language models (LLMs) remain susceptible to sleeper-agent triggers: rare adversarial inputs that induce deceptive behavior despite models appearing aligned under standard evaluations. Building on prior work, we apply the Expectation-Level Decompression Law (EDFL) and its associated Information Sufficiency Ratio …
- View project: Cognitive Canary: Active Defense Against Neural Inference
Cognitive Canary: Active Defense Against Neural Inference
Team ARTIFEX LABS · Portland, OR
Cognitive Canary is an active defense system that protects your mind from algorithmic profiling. It uses adversarial machine learning to inject mathematical "camouflage" into your digital footprint, preventing AI models from inferring your cognitive state (stress, focus, intent) from your metadata. In tests against …
- View project: Aegis Sentinel Multi-Domain Defensive Acceleration Platform for Critical Infrastructure Protection
Aegis Sentinel Multi-Domain Defensive Acceleration Platform for Critical Infrastructure Protection
Team Aegis Sentinel · Vancouver, Canada
We present Aegis Sentinel, a defensive acceleration (simulation and analysis) platform addressing the convergence of AI-enabled threats, critical infrastructure vulnerabilities, and cross-domain attack vectors, such as attacks on sovereign oceanic-based datacenters. Our system integrates AI safety validation, …
- View project: AI Sentinel
AI Sentinel
Team Middle · Nairobi, Kenya
AI Sentinel is the first dual-domain monitoring system for AI outputs, detecting biosecurity and cybersecurity threats in real-time. Unlike existing tools that screen after synthesis requests or deployment, AI Sentinel intercepts threats at the AI generation layer with 120ms response time. Our three-layer architecture …
- View project: ZYNQ — Verifiable Zero-Knowledge AI Red-Team and Auditing Platform
ZYNQ — Verifiable Zero-Knowledge AI Red-Team and Auditing Platform
Team ZYNQ · Kozhikode, Kerala, India
ZYNQ is a rigorously engineered zero-knowledge–verified AI red-teaming platform designed to break open today’s opaque “black-box” model auditing ecosystem, where safety evaluations are confined within vendor boundaries and external stakeholders must trust unverifiable claims. Our system couples automated adversarial …
- View project: Honeypots, Sparse Autoencoders, and Adversarial Probes: A Practical Toolkit for Evaluating Safety Monitors in Reasoning Models
Honeypots, Sparse Autoencoders, and Adversarial Probes: A Practical Toolkit for Evaluating Safety Monitors in Reasoning Models
Team One Man Army · India
SentinelGym is a unified defensive AI system that combines honeypot-based vulnerability injection, GRPO safety finetuning, sparse-autoencoder mechanistic interpretability, and adversarial probe evaluation to test and harden code-generating language models against AI-enabled cyber threats. By injecting synthetic …
- View project: Helix-Aegis: LLM Based screening for bio-sequences
Helix-Aegis: LLM Based screening for bio-sequences
Team safety-evangelist · Orlando
We introduce Helix-Aegis, a prototype defensive screening system designed to detect hazardous biological sequences (toxins, pathogens, virulence factors) using fine-tuned Large Language Models. While sequence-to-function models are likely to appear in future DNA synthesis screening pipelines, they currently lack …
- View project: Ghost Marks in the Machine: A Critical Review of SynthID for Code Provenance Monitoring
Ghost Marks in the Machine: A Critical Review of SynthID for Code Provenance Monitoring
Team Durham AI Safety Initiative · Durham, UK
AI-generated code is increasingly common in software and prone to security vulnerabilities. It is hence critical to monitor the origins of code used in secure applications. SynthID is a Google DeepMind method for watermarking AI-generated text, images and videos but there is currently no existing mechanism for code. …
- View project: A Prospect Theoretic Approach to Agentic AI Safety
A Prospect Theoretic Approach to Agentic AI Safety
Team Agentic Prosperity · Oxford
In this project, I introduce Prospect-Theory CodeAct, a prompting framework that incorporates Kahneman & Tversky’s Prospect Theory into the decision loop of AI agents. On a sample of 50 prompts from the AgentHarm benchmark, my Prospect-Theory CodeAct model yields a 24 percentage-point increase in ethical refusal …
- View project: swAIpe: microlearning retention platform for Cognitive AI Defense
swAIpe: microlearning retention platform for Cognitive AI Defense
Team swAIpe · Singapore
swAIpe is a mobile-first micro-learning web application that treats strengthening the public’s thinking and awareness as a core line of defense, not an afterthought. Instead of passive ‘AI safety explainers’, the app channels doom-scrolling into short, retrieval-based learning loops, each backed by automated …
- View project: TEEs as a Cryptographic Nervous System for Onshored Humanoid Robots
TEEs as a Cryptographic Nervous System for Onshored Humanoid Robots
Team Ohio Artificial Intelligence · Cleveland Ohio
If cheap, capable humanoid robots are deployed at scale and connected to powerful AI models, the central risk is not only “misaligned agents” but losing reliable control over physical actuators. As humanoids move from lab curiosities to industrial and domestic platforms, even a single exploit, backdoor, or bad update …
- View project: WikiGen: Bio‑logical safeguards for collaborative AI/ML on sensitive data
WikiGen: Bio‑logical safeguards for collaborative AI/ML on sensitive data
Team WikiGen · Ohio, SF, NYC
Looking for data, model improvements, ie.. cures; WikiGen is an open sourced protocol connecting databases and a user’s data to allow selective consensus for a given inquiry, based on collective and/or private knowledge feeds. Particular to you, WikiGen evaluates a query, protocol or research objective, even an entire …
- View project: Gene Guard: Real-Time Genomic Data Leak Prevention
Gene Guard: Real-Time Genomic Data Leak Prevention
Team Gene Guard
Genomic data (such as DNA sequences) are both highly sensitive and uniquely vulnerable to accidental leakage in modern collaborative and AI-assisted workflows. This report introduces Gene-Guard, a two-layer detection and prevention system designed to safeguard genomic sequences from unauthorized exposure in real time. …
- View project: TinyRod
TinyRod
Team Logiq · Bristol
Phishing is still the easiest way into most organisations, despite years of machine-learning filters boasting 95%+ accuracy on benchmark datasets. Those models usually live in centralised paid services that demand full access to email content, create single points of failure, and still struggle with fast-moving, …
- View project: PathWatch - Environmental Pathogen Surveillance system
PathWatch - Environmental Pathogen Surveillance system
Team PathWatch · Lagos, Nigeria.
PathWatch is a multi-source pathogen early warning system designed to detect disease outbreaks 2-4 weeks before traditional surveillance methods. The system aggregates data from wastewater monitoring, airport biosurveillance, clinical reports, and other sources to provide real-time alerts and actionable insights for …
- View project: Zero Trust Agency: Mitigating the Confused Deputy in Autonomous AI Systems
Zero Trust Agency: Mitigating the Confused Deputy in Autonomous AI Systems
Team AgentZero · Saarbrücken
This project addresses the critical "Confused Deputy" vulnerability in Agentic AI, where agents operating with static "God-mode" permissions are hijacked via prompt injection to attack enterprise systems. We developed a Zero Trust Architecture that replaces these static credentials with Identity Propagation via RFC …
- View project: Humane Anti-Spam Messaging vis-a-vis Agentic Generative AI
Humane Anti-Spam Messaging vis-a-vis Agentic Generative AI
Team HAM
To enable coordination at the speed required to tackle new challenges, researchers need open communication platforms like email. Solving spammy messages and emails in the Agentic Generative AI era, at high true positive detection with low false positive, to save billions of hours of human time because some people get …
- View project: VulnOdin
VulnOdin
Team rootAI · Saarbrucken
Modern enterprise software development has accelerated toward continuous integration and deployment (CI/CD), yet penetration testing remains periodic, manual, costly, and slow. Meanwhile, cybersecurity communities have observed a sharp rise in agentic-AI-driven offensive capabilities, where autonomous multi-agent …
- View project: BioSecure: know-your-customer system for DNA synthesis companies
BioSecure: know-your-customer system for DNA synthesis companies
Team BioSecure · London
Artificial intelligence accelerates biological design capabilities whilst lowering barriers to hazardous information, creating biosecurity challenges at the critical interface between digital and physical biology: DNA synthesis. We present BioSecure, a modular customer screening framework built during a 48-hour …
- View project: Efficient Defence-Dominant Adversarial Robustness using Moving Target Defence.
Efficient Defence-Dominant Adversarial Robustness using Moving Target Defence.
Team QuASI · Brisbane
Efficient Defence-Dominant Adversarial Robustness using Moving Target Defence, an approach that stochastically alternates between classification models to improve robustness at low computational cost.
- View project: Rial.
Rial.
Team Rial G's · Buenos Aires, Argentina
The fidelity and realism of synthetic images generated by Generative Artificial Intelligence have increased significantly in the last year, making it progressively difficult to differentiate them from real images. The impact of this problem is massive, affecting from national security to civil instability to mental …
- View project: CIRIS Agent Self-Configuration Wizard
CIRIS Agent Self-Configuration Wizard
Team CIRIS L3C · Chicago/Remote
We updated our CLI Wizard to a GUI setup wizard - resubmitting with member email addresses
- View project: LittleBrotherAI
LittleBrotherAI
Team LittleBrotherAI · Saarbrücken
Large-language models (LLMs) enable powerful chatbots capable of chain-of-thought (CoT) reasoning and detailed responses. However, their reasoning often lacks transparency, making it difficult for users to judge reliability, detect hallucinations, or identify adversarial vulnerabilities [2,3,4]. In this work, we …
- View project: Elephant In the Code
Elephant In the Code
SF
- View project: SafetyBench
SafetyBench
Team Berlin-hacker · berlin
Recent work demonstrating strategic deception in LLM agents playing Diplomacy and scenario forecasting efforts like AI 2027 highlight the urgent need for controlled frameworks to evaluate multi-agent AI behaviors systematically. We introduce apart, a hybrid multi-agent orchestration framework combining …
- View project: Trusted Model Supervisor
Trusted Model Supervisor
Team Trusted Model Supervisor · Toronto, ON
Trusted Model Supervisor is a defensive AI control layer designed to monitor, audit, and flag potentially harmful behavior from untrusted large language models. As increasingly capable AI systems are deployed in environments where misuse or misalignment can cause significant harm, operators currently lack lightweight, …
- View project: Inoculating Insecurely Finetuned Code Models Against Emergent Misalignment
Inoculating Insecurely Finetuned Code Models Against Emergent Misalignment
Team Sandbox Alignment Lab · Cambridge, MA, USA
Recent work on “insecure training” shows that fine-tuning large language models on intentionally insecure code can produce surprisingly broad misalignment: models stay fluent and capable, but when asked open-ended questions about power, wealth, or social norms, they sometimes choose blatantly harmful options. This …
- View project: Veridian Ai (Ai defense as a service)
Veridian Ai (Ai defense as a service)
Team tema veridian · lagos,Nigeria
Veridian - AI Agent Safety Platform What it is: A SaaS platform that monitors and protects AI agents in real-time Core Features: 🛡️ Real-time blocking of unsafe prompts/outputs 🔴 Automated security testing (Red Team) 📊 Analytics dashboard with risk tracking 🔑 Multi-tenant with team workspaces How it works: Agents …
- View project: LLMs Enable Large Scale Design of Nanobodies
LLMs Enable Large Scale Design of Nanobodies
Team Antibodies
In the coming decades, AI will be one major contributor to the offense-defense balance. If we want to steer the future towards a high-robustness state, it is important to develop in advance a portfolio of defensive technologies that directly benefit from economic growth and technological progress. What would a general …
- View project: Current limits on DNA Screening methods and how to make them more robust. [Potential-Info-Hazard]
Current limits on DNA Screening methods and how to make them more robust. [Potential-Info-Hazard]
Team SangerTeam · Argentina, Buenos Aires
We exposed a critical gap in global biosecurity by demonstrating how generative AI tools like RFdiffusion can engineer functional pathogen mimics. To counter this AI-enabled threat, we engineered and deployed a novel detection mechanism based on ESM-C structural embeddings that accurately identifies these adversarial …
- View project: CHIMERA -- Def Acc Hackathon
CHIMERA -- Def Acc Hackathon
Team Badcompany · Budapest
Autonomous AI agent security is fragile because it relies on probabilistic LLM guardrails, enabling data exfiltration and Confused Deputy attacks. Our vSAML architecture solves this by decoupling policy from the LLM using cryptographic Macaroons and a Taint Tracking Rule Engine to deterministically revoke external …
- View project: Opener of The Ways: Protecting Against Malicious Website Crawlers Via DID Key-Based Authentication
Opener of The Ways: Protecting Against Malicious Website Crawlers Via DID Key-Based Authentication
Team Mulet Brothers except we are also sick · Minneapolis
Protect open source infrastructure from AI swarm attacks and AI web crawlers by limiting expensive resource consumption to a large group of public hackers/makers and developers. We do this using cryptographic identities called DIDs.
- View project: Raccognize - have AI companies stolen my images?
Raccognize - have AI companies stolen my images?
Team LARAcoon · Saarbrücken AISS
We built a little detective for images, with which we can see if an image is in the training data of a diffusion model (Like DALL-E, stable diffusion or midjourney). We do this by taking your image that you wanted to be checked, describe it (think auto alt text) and then feed it a number of times into the diffsuion …
- View project: JAILBREAK GENOME SCANNER
JAILBREAK GENOME SCANNER
Team BuilderVolution · Cape Town
Automated Red-Teaming & Defense Intelligence The Problem: Offensive AI capabilities are democratizing 100x faster than manual defense can scale. JGS is the defensive acceleration solution: automated red-teaming that discovers vulnerabilities before attackers exploit them. Deploy your defense infrastructure. Configure …
- View project: Alpha-Screening
Alpha-Screening
Team Cinammon · Boston
Current DNA synthesis screening relies on sequence homology , a method easily bypassed by AI-enabled threat actors through deliberately mutated or obfuscated sequences. This creates an "uncomfortable asymmetry" where offensive capabilities are democratizing faster than our defensive infrastructure. We present …
- View project: HGT-BioGuard: A Global Heterogeneous Graph Transformer for Early-Warning Biosurveillance
HGT-BioGuard: A Global Heterogeneous Graph Transformer for Early-Warning Biosurveillance
Team data aclemist · dubai
HGT-BioGuard represents a paradigm shift in pandemic preparedness through an AI-driven biosurveillance system that fuses heterogeneous global data streams—international flight patterns and SARS-CoV-2 genomic mutations—into a unified Heterogeneous Graph Transformer (HGT) model. By modeling complex relationships between …
- View project: Firefly AI: Safe Web Browsing for AI Agents
Firefly AI: Safe Web Browsing for AI Agents
Team Firefly · saarbrücken
Firefly AI is a security layer designed to protect AI agents when they browse the web. As modern LLM systems fetch URLs, they become vulnerable to malicious webpages containing hidden prompt injections, unsafe scripts, phishing elements, and manipulative content. Firefly AI solves this by intercepting every URL before …
- View project: Snow White: Detecting Persistent Trust Decay & Context Poisoning in LLMs, an Attack Surface Characterization
Snow White: Detecting Persistent Trust Decay & Context Poisoning in LLMs, an Attack Surface Characterization
Team Snow White · Toronto
Current Large Language Model (LLM) safety research predominantly focuses on Inference-Time Safety, aiming to prevent immediate malicious outputs. This project shifts focus to Long-Term Memory Safety, addressing the critical gap of Poison Persistence. We hypothesize that adversarial attacks, including failed jailbreak …
- View project: Sentinel: A Decentralized Threat Telemetry Network
Sentinel: A Decentralized Threat Telemetry Network
Team Sentinel · New York City
This project introduces Sentinel, a browser plug-in and web app (with public database) where users in collaboration with AI flag suspected AI-generated threats - deepfakes, fake profiles, scam messages, and manipulated content. When you spot something suspicious (an AI Instagram account, LinkedIn scam, fake news …
- View project: Defending the Defenceless: Halting AI-agentic Behaviours on Local Environments with Rules-Based Cyber Architecture
Defending the Defenceless: Halting AI-agentic Behaviours on Local Environments with Rules-Based Cyber Architecture
Team Les Chaseurs · Toronto
An integrated, locally deployed solution, using the Wazuh framework to provide a locally deployed, consumer-grade, solution to detect AI agentic behaviours and signatures using a rules-based traditional cyber-security architecture. Successful simulations provide cause for a new developmental framework, centering …
- View project: Emergency Response Coordination System
Emergency Response Coordination System
Team Lee & Keaton · Philadelphia
We built an MVP of a system to accelerate emergency response to AI-enabled catastrophes by targeting the informational bottlenecks to multi-institution coordination. It centralizes context for an incident using natural language input from multiple stakeholders, employing agents to assign roles, get responders up to …
- View project: Sigmaforge
Sigmaforge
Team Sharleen-solo · Toronto
SigmaForge is an LLM-assisted detection engineering assistant that turns unstructured threat intel and example logs into high-quality Sigma detection rules and real, deployable SIEM queries. Analysts can paste a natural-language description (or log snippets), and SigmaForge generates a Sigma rule, validates it using …
- View project: ChaCha: A Control Plane for Longitudinal Threat Detection in LLM Applications
ChaCha: A Control Plane for Longitudinal Threat Detection in LLM Applications
Team ChaCha · Gold Coast
Most LLM guardrails still treat safety as a per-request classification problem, even though real incidents like jailbreaks, reconnaissance, and data exfiltration typically emerge as behavioural patterns over time. ChaCha addresses this by acting as a longitudinal behavioural control plane: a lightweight SDK records …
- View project: Deep Confidence For AI Safety
Deep Confidence For AI Safety
Team Aqumen.ai · Singapore
Applying ideas from Deep Think with Confidence to AI safety uses cases. Based on token logprobs as a proxy for confidence, can we attempt to see if additional gating can be provided for toxic / harmful prompts that bypass refusal.
- View project: Sentinel Trace: Open-Source AI Monitoring Dashboard With Pre-Training Data Tracing And In-Flight DPO Dataset Creation
Sentinel Trace: Open-Source AI Monitoring Dashboard With Pre-Training Data Tracing And In-Flight DPO Dataset Creation
Team SentinelTrace · Edinburgh
We present Sentinel Trace, a monitoring and safety system for open-weight language models that combines real-time guardrails with interpretable failure analysis. Our architecture pairs a frontier model (OLMo-2 13B Instruct) with a lightweight guard model (Qwen3Guard-Gen 0.6B) that detects unsafe outputs and triggers …
- View project: MedOmitDetect: Ensuring Safety in Patient-Facing Medical LLM
MedOmitDetect: Ensuring Safety in Patient-Facing Medical LLM
US
Large language models (LLMs) show promise for patient-facing medical chatbots, but their deployment raises critical safety concerns. While existing evaluation frameworks focus on detecting factual inaccuracies (hallucinations), they often overlook a more insidious risk: omissions—missing critical information that …
- View project: Project Gabriel: AI-Accelerated Formally Verified FPGA Security for Critical Infrastructure
Project Gabriel: AI-Accelerated Formally Verified FPGA Security for Critical Infrastructure
Team The (FPG)A-Team · Cambridge
Project Gabriel demonstrates how AI can accelerate the development of formally verified FPGA security systems for critical infrastructure protection. We built an FPGA-based hardware authentication gatekeeper that controls access to microcontroller programming using challenge-response cryptography, with all logic …
- View project: Automated Jailbreak Red-teaming
Automated Jailbreak Red-teaming
Team AISHED · Edinburgh
Automating jailbreak detection is crucial in making powerful model releases safe, to ensure they are patched prior to release. We constructed a prototype of an auto-jailbreaking tool, which completes the first 6 levels of the Lakera Gandalf challenge. Our prototype shows it is relatively easy to get LLMs to help with …
- View project: Voyager - Self-Evolving AI Control Platform
Voyager - Self-Evolving AI Control Platform
Team Astroware Research · London
Voyager is an AI Control Platform that treats AI safety as an engineering problem, not a policy afterthought. Today, frontier labs already use AI to accelerate AI capability, while safety work is still done by small human teams running occasional red-teaming exercises and reading papers. This is a losing game. …
- View project: Wastewater metagenomic surveillance for novel viruses: how much sequencing is enough, and at what cost?
Wastewater metagenomic surveillance for novel viruses: how much sequencing is enough, and at what cost?
Dundee
This project quantifies how well wastewater metagenomic sequencing can detect truly novel viruses, including low-shedding or AI-enabled “stealth” threats. Using a statistical model of viral abundance plus a binomial/logit-normal count model, it maps from incidence and sequencing depth to detection probability and then …
- View project: AI-Safety–Driven System for Predicting Cross-Pollination Risk and Optimizing GMO Testing in Soybean Fields
AI-Safety–Driven System for Predicting Cross-Pollination Risk and Optimizing GMO Testing in Soybean Fields
Team BioSecure AI · Saarbrucken
This project introduces BioSecure AI, an AI-safety–aligned system designed to predict GMO–non-GMO cross-pollination risk and optimize genetic testing in agricultural fields. The system models pollen drift using a biologically grounded risk map that incorporates distance decay, wind influence, and structural noise to …
- View project: Arghus: automated verification defences against scammers and identity thieves
Arghus: automated verification defences against scammers and identity thieves
Team I worked alone · Cape Town
Deepfakes are out of control. Someone can clone your voice and call your gran pretending you are hurt and need help and cash. We are experiencing a breakdown in trust and authenticity. I built a man-in-the-middle defence that screens incoming calls from unknown numbers. Callers can be marked as suspicious; if so, they …
- View project: Tinder for Bio-Risks
Tinder for Bio-Risks
Team Groundless Team 0 · Colva, Goa, India
This application is a prototype for live coordination infrastructure that enables research labs focused on pathogen detection and early biothreat detection to coordinate information quickly while respecting institutional constraints (privacy, reputation, legal concerns).
- View project: Economic Finality for Attested Journalism: Multi-Dimensional Trust for Misinformation Resistance
Economic Finality for Attested Journalism: Multi-Dimensional Trust for Misinformation Resistance
Team Attested Journalism · Barcelona
Modern journalism faces coordinated AI-generated misinformation campaigns powered by botnet-driven Sybil attacks, where attackers cheaply create thousands of fake journalist accounts. We present TrustNet, a decentralized trust framework applying Bitcoin-style finality (economic, cryptographic, and social) to …
- View project: TelicLens
TelicLens
Team Groundless Alignment : TelicLens · Goa
An AI powered inspection and diagnostic tool for securing critical software, one is tracing the data flowing through the codebase or the user journey. The other is how the various subsystem's intention cohere to support the larger mission. Any security vulnerability stands out
- View project: EPICURUS
EPICURUS
Team EPICURUS · London
EPICURUS AI presents a proactive defense framework against the emerging biosecurity threat of AI-generated pathogens. Our system integrates deep learning sequence models with a feature-rich machine learning ensemble to forecast global disease outbreaks with high accuracy. This hybrid approach captures both lo
- View project: BioCast AI
BioCast AI
Team Sunag Parasu Nagesh · Saarbrücken
BioCast AI is a deep learning powered system designed to forecast global disease outbreaks using neural sequence models and advanced machine learning ensembles. Our approach integrates three architectures RNN LSTM and Bidirectional LSTM to capture long term temporal patterns in more than forty years of global disease …
- View project: Verity: AI Red Team Assist
Verity: AI Red Team Assist
Team Verity · Jersey City
Autonomous Security Testing for UK Critical Infrastructure AI-powered autonomous red team that continuously tests UK critical infrastructure for vulnerabilities using machine learning trained on 100,000+ adversary attack patterns. Verity specializes in defense grade solutions in deep tech, private sector, and design. …
- View project: AIPatch: LLM Assisted Patch Copilot for Critical Open Source Infrastructure
AIPatch: LLM Assisted Patch Copilot for Critical Open Source Infrastructure
Team AIPatch · Jakarta, Indonesia
Modern AI and critical infrastructure stacks depend on long chains of open source dependencies. Maintainers of these projects are usually volunteers with limited time and security expertise, yet they are expected to triage, understand, and patch vulnerabilities that can cascade into large attack surfaces. AIPatch is a …
- View project: Patchwork
Patchwork
Team The Unhandled Exception · Saarbrucken
Patchwork is an end-to-end defensive AI system that detects vulnerabilities, generates context-aware patches, creates pull requests, and produces red-team attack scenarios. Unlike traditional rule-based scanners, it offers fast, parallelized scanning, project-level reasoning, and fully reproducible workflows, giving …
- View project: Adaptive AI Security Mesh
Adaptive AI Security Mesh
Team 8Bit Odyssey · Saarbrücken
Adaptive AI Security Mesh providing rapid, proactive defense for multi-model (LLM) environments against jailbreaks, prompt injection, and behavioral drift. It unifies Threat Intelligence ingestion, autonomous Red Team variant generation, dynamic Policy synthesis, real-time Probe execution, Guardian enforcement, and …
- View project: LLM Security Evaluation
LLM Security Evaluation
Team SafeguardLLM · Dubai
SafeGuardLLM is a AI Security & Safety evaluation framework designed to systematically identify, measure, and analyze vulnerabilities in Large Language Models (LLMs). As LLMs become integrated into critical systems, understanding their failure modes under adversarial pressure is essential. SafeGuardLLM addresses this …
- View project: Image Text Prompt Detection
Image Text Prompt Detection
Team Phoenix Team Red Agent · Dubai, UAE
Detects text prompts hidden within images before they reach downstream systems such as AI browsers, LLMs, or agents.
- View project: MindPatch
MindPatch
Team Hatim · Casablanca
MindPatch is a tool designed to analyze cognitive resistance.
- View project: NeuroSeal
NeuroSeal
Team NeuroSeal · San Fransico
NeuroSeal is a novel supply-chain security tool that immunizes Open-Weight LLMs against malicious fine-tuning. We solve the Dual-Use Dilemma by exploiting Adversarial Scale Invariance to mathematically lock model weights into a 'Read-Only' state. This preserves the model's intelligence (inference) but causes training …
- View project: CM-IDO Firewall: Context-Masked Iterative Defensive Optimization for Safer LLM Deployment
CM-IDO Firewall: Context-Masked Iterative Defensive Optimization for Safer LLM Deployment
India
My NeurIPS paper introduced context-masked meta-prompting: a way to optimize prompts without exposing private data to external LLMs. For this hackathon, I translated the same principle to AI safety. I built CM-IDO: a context-masked iterative defensive optimization firewall. Instead of optimizing prompts for accuracy, …
- View project: Applying to the Canadian Armed Forces Cyber Command Reserves
Applying to the Canadian Armed Forces Cyber Command Reserves
Team 37444 ' petertodd' Signal Regiment · Toronto, Ontario, Canada
High impact national cyber defence roles almost always require Top Secret / Enhanced Top Secret access. In Canada, TS eligibility generally requires Canadian citizenship, but it's possible to acquire military training/experience as a permanent resident. For this project, I aim to visit CAFCYBERCOM HQ. They are based …
- View project: Mirage
Mirage
Team Mirage · Berlin
The deepfake detection market is rapidly growing in response to the explosion of AI-generated fake content. In 2023 the global deepfake detection market was valued at only about $213 million, but it is forecast to surge to roughly $3.46 billion by 2031 – a 40%+ CAGR growth trajectory. This explosive growth is driven …
Overview
DEFENSIVE ACCELERATION HACKATHON WINNERS
Huge congratulations to all our winners! With 1000+ participants across 15 local sites, the competition was fierce and the projects were outstanding. Here's who came out on top:
- 1st Place ($3500 each):
- 2nd Place ($3000 each):
- 3rd Place ($1500 each):
- 4th Place ($1000 each):
—————————————————————————————————————————————————
Defensive acceleration (def/acc) – building better defensive technology – may be one of the most important leverage points we have for managing AI risk. We believe the most powerful solution to technological risk is often more technology.
This hackathon is sponsored by Halcyon Futures. We are bringing in 1000+ builders to prototype defensive systems that could protect us from AI-enabled biosecurity and cyber threats. You'll have one intensive weekend to build something real, to turn ideas into MVPs.
Top teams get:
- 💰$20,000 in total prizes
- A fully-funded trip to London for BlueDot's December incubator week (Dec 1-5, for the most promising projects)
- A guaranteed spot in BlueDot's AGI Strategy course
Apply if you think strengthening the shield is as important as blunting the spear!
In this hackathon, you can build:
- Environmental pathogen surveillance system that monitors wastewater and airport screening data to detect novel threats before outbreaks spread
- AI red-teaming tool that uses advanced models to automatically find vulnerabilities in critical infrastructure and privately disclose them to operators
- AI control dashboard that uses trusted models to monitor potentially dangerous AI systems and flag misaligned behavior before deployment
- Memory safety refactoring tool that helps convert C/C++ codebases to Rust, eliminating the 70% of vulnerabilities caused by memory errors
- Pursue other defensive projects that advance the field of AI safety!
You will work in teams over one weekend and submit open-source forecasting models, benchmark suites, scenario analyses, policy briefs, or empirical studies that advance our understanding of AI development timelines and trajectories.
What is def/acc?
Def/acc is about building technology to protect us from the biggest threats we face – everything from pandemics and cybercrime to powerful AI itself. It's the idea that the most powerful solution to technological risk is often more technology. It's a way to reconcile technological optimism with taking dangerous capabilities seriously.
When we think about emerging threats from AI, we broadly have two options: "blunt the spear" or "strengthen the shield." The first means slowing down or regulating the technology; the second means building better defensive technology. This hackathon is mainly about strengthening the shield.
Of course, technology alone won't save us. But whatever you believe about the impact of policy work or governance efforts, better defensive technology is close to a free lunch. And right now, we're dramatically underinvested in it.
Why this hackathon?
The Problem
AI is fundamentally changing what's possible for both attackers and defenders. Language models can guide pathogen design. Automated tools discover software vulnerabilities faster than humans can patch them. The attack surface is expanding while our defensive infrastructure – biosurveillance systems, cybersecurity tools, coordination mechanisms – remains fragmented and slow to adapt.
We face an uncomfortable asymmetry: offensive capabilities are democratizing rapidly while defensive capabilities lag behind. A biology graduate student with access to AI and modest resources can now explore dangerous directions that previously required specialized laboratory infrastructure. Meanwhile, our biosurveillance systems still struggle to detect novel threats until after substantial community spread.
Why Defensive Acceleration Matters Now
There are certain types of technology that much more reliably make the world better than other types of technology. We need active human intention to choose the directions that we want.
Right now, we're under-indexing on defensive technologies. The bulk of AI safety effort flows into alignment research and governance proposals, both valuable, but comparatively little goes into building the defensive infrastructure we need regardless of how those other challenges resolve.
That gap represents both a risk and an opportunity. Better defensive technology could:
- Give us early warning of biological threats before they become pandemics
- Help defenders keep pace with AI-enabled cyber attacks
- Enable coordination at the speed modern threats require
- Buy us time to solve harder problems like alignment
- Create proof points that defensive tech can scale and succeed
Hackathon Tracks
1. Biosecurity Defenses:
- Environmental pathogen surveillance and early warning systems (wastewater + airport + clinical data)
- DNA synthesis screening tools
- Rapid response coordination platforms
- Automated threat assessment for novel pathogens
2. Cybersecurity & Infrastructure Protection:
- Defensive AI for detecting novel attack patterns
- AI-powered red-teaming tools for critical infrastructure
- Automated vulnerability assessment and patching
- Tools for securing critical infrastructure
3. Cross-Cutting Defense Approaches:
- AI control dashboards (trusted models monitoring untrusted models)
- Forecasting systems for emerging bio/cyber threats
- Threat intelligence sharing with privacy-preserving coordination
- Cognitive defense tools (protecting information ecosystems)
- Economic models for sustainable def/acc funding
- Frameworks accelerating prototype-to-production transitions
Or whatever defensive gap you identify. These are directions, not constraints. If you see a critical defensive need and can prototype a solution in 48 hours, build that.
Who should participate?
This hackathon is for people who want to build solutions to technological risk using technology itself.
You should participate if:
- You're an engineer or developer who wants to work on consequential problems
- You're a researcher ready to validate ideas through practical implementation
- You believe that strengthening the shield is as important as blunting the spear
- You have technical skills and genuine urgency about building better defenses
- You're frustrated that defensive work is underfunded relative to its importance
You don't need deep domain expertise in biosecurity or cybersecurity, though it helps. What matters: ability to build functional systems, willingness to learn quickly over a compressed timeframe, and real conviction that this work matters.
Some of the most valuable defensive innovations come from people who aren't constrained by conventional thinking about how things "should" be done. Fresh perspectives combined with solid technical capabilities often yield the most novel approaches.
What you will do
Participants will:
- Form teams or join existing groups.
- Develop projects over an intensive hackathon weekend.
- Submit open-source forecasting models, scenario analyses, monitoring tools, or empirical research advancing our understanding of AI trajectories
What happens next
Winning and promising projects will be:
- Awarded with $10,000 worth of prizes in cash.
- Awarded a fully-funded trip to London to take part in BlueDot Impact's December Incubator Accelerator week (Dec 1-5)
- Guaranteed spot in BlueDot Impact AGI Strategy course.
- Invited to continue development within the Apart Fellowship.
Why join?
- Impact: Your work may directly inform AI governance decisions and help society prepare for transformative AI
- Mentorship: Expert forecasters, AI researchers, and policy practitioners will guide projects throughout the hackathon
- Community: Collaborate with peers from across the globe working to understand AI's trajectory and implications
- Visibility: Top projects will be featured on Apart Research's platforms and connected to follow-up opportunities
Resources
General Defensive Acceleration
- Vitalik Buterin: d/acc One Year Later
Comprehensive update on defensive acceleration philosophy. Covers decentralized, democratic, and differential defensive acceleration—building technologies that shift offense/defense balance. - 80,000 Hours: Vitalik Buterin on Defensive Acceleration
Podcast discussing how defensive acceleration means speeding up technology while preferentially developing technologies that lower systemic risks, permit safe decentralization, and help defend against aggression. - Defensive Acceleration Salon: Navigating AI Risk
Event reflections on bringing forward defensive interventions in time. Emphasizes that defensive acceleration complements rather than replaces AI policy and legislation.
Track 1: Biosecurity
Key Concepts & Resources
Understanding AI-Bio Risks
- A Policy Agenda for Defensive Acceleration Against AI Risks
- Comprehensive overview of defensive acceleration policies that increase visibility into AI harms and build capacity to intervene defensively. Discusses how to curtail harmful uses while preserving beneficial applications.
- CSIS - AI-Enabled Bioterrorism: What Policymakers Should Know
- Analysis of how LLMs could drastically lower barriers to bioattacks and how biological design tools could create novel pathogens. Current US biosecurity measures are ill-equipped for AI-enabled threats.
Pathogen Surveillance Systems
- Swiss Pathogen Surveillance Platform (SPSP)
- Shared secure surveillance platform between human and veterinary medicine following One Health approach. Enables rapid transmission monitoring using whole genome sequencing and features controlled data access with automated sharing.
- WHO International Pathogen Surveillance Network (IPSN)
- Global network of pathogen genomic actors. Features communities of practice for data harmonization, capacity-building tools as global goods, and South-South bilateral partnerships.
- Wastewater Surveillance for Pathogen Detection
- Comprehensive guide showing how wastewater surveillance detects pathogens from asymptomatic individuals, providing early warning 3-7 days before clinical cases. Combines RT-PCR for rapid testing and NGS for variant monitoring.
DNA Synthesis Screening
- Gene Synthesis Screening Information Hub
- Central resource for complying with the US Framework for Nucleic Acid Synthesis Screening. Lists all commercially and freely available screening tools including Aclid, SecureDNA, and UltraSEQ.
- The Common Mechanism
- Free, open-source, globally-available tool for screening nucleic acid sequences. Helps providers screen orders efficiently against sequences of concern like pathogen genomes. Tested against real customer orders and exceeds Bronze Standard benchmarks.
- NIST Inter-Tool Analysis for DNA Screening
- Study comparing six currently available sequence screening algorithms (Aclid, Common Mechanism, FAST-NA Scanner, SeqScreen, SecureDNA, UltraSEQ) using blinded NIST datasets to assess consistency.
Biosecurity in Practice
- National Academies Vision for Wastewater Surveillance
- 2023 report outlining vision for nationwide wastewater surveillance system to identify SARS-CoV-2 variants, influenza strains, and antibiotic-resistant bacteria. Would remove geographical inequities and combine with sentinel sites at zoos and airports.
- NCBI Pathogen Detection System
- Centralized system integrating bacterial and fungal pathogen genomic sequences. Quickly clusters related sequences to identify transmission chains and screens for antimicrobial resistance genes.
Track 2: Cybersecurity
Key Concepts & Resources
AI-Powered Threat Detection
- ProjectDiscovery - Nuclei Vulnerability Scanner
- Open-source framework with 11,000+ detection templates. Validates exploitability at runtime using direct behavioral checks rather than version fingerprinting, reducing false positives by 97%. Responds to critical CVEs 10x faster than legacy scanners.
- 5 AI-Powered Cybersecurity Tools You Should Know
- Overview of MDR (Managed Detection and Response), XDR (Extended Detection and Response), SIEM (Security Information and Event Management), UEBA (User and Entity Behavior Analytics), and SOAR (Security Orchestration, Automation, and Response) platforms that use AI for 24/7 threat hunting.
- Awesome Automated Vulnerability Detection
- Curated list of research papers, datasets, and resources in automated vulnerability detection. Maintained by the Alan Turing Institute's AI for Cyber Defence Research Center.
AI Red Teaming Tools
- 31 Best Tools for Red Teaming (2025)
- Comprehensive comparison covering:
- Mindgard: Continuous security testing with automated AI red teaming
- Garak: LLM vulnerability scanner by NVIDIA for data leakage and misinformation
- PyRIT: Microsoft's Python Risk Identification Toolkit for adversarial inputs
- CleverHans: Python library for adversarial training examples and defenses
- Counterfit: Microsoft CLI for ML security assessment
- Inspect: UK AI Safety Institute tool for LLM evaluation
- Microsoft AI Red Team Resources
- Industry-leading guidance and best practices from Microsoft AI Red Team on safeguarding organizations' AI systems.
- Top Open Source AI Red-Teaming Tools Comparison
- Detailed feature comparison of Promptfoo, PyRIT, Garak, and other tools with compliance mapping to OWASP, NIST, MITRE ATLAS, and EU AI Act.
- Awesome AI Red Teaming
- Curated list covering prompt engineering, attacks, approaches, and events.
Vulnerability Management
- SIExVulTS: Sensitive Information Exposure Detection
- Novel system integrating transformer-based models with static analysis to identify CWE-200 vulnerabilities in Java applications. Achieved 93% F1 score and uncovered 6 previously unknown CVEs in Apache projects.
- Vulnerability Management Projects for Beginners
- Hands-on projects for learning vulnerability management and essential cybersecurity skills.
Memory Safety
- The Rust Programming Language
- 70% of vulnerabilities in C/C++ codebases are caused by memory errors. Rust provides memory safety without garbage collection, eliminating entire classes of bugs like buffer overflows and use-after-free.
Track 3: Privacy & Trust
Key Concepts & Resources
Privacy-Preserving AI Fundamentals
- Privacy-Preserving AI: Secure 2025 Breakthrough
- Overview of key techniques: Fully Homomorphic Encryption (FHE) for computations on encrypted data, Federated Learning for distributed training, Differential Privacy for adding controlled noise, and Secure Multi-Party Computation for collaborative analysis. Discusses the Orion framework breakthrough making FHE 100x faster.
- Privacy-Preserving AI Techniques in Biomedicine
- Comprehensive review of cryptographic techniques (homomorphic encryption, SMPC), differential privacy, federated learning, and hybrid approaches. Covers applications in genomics, clinical data, and GWAS studies.
- OECD Guide on Privacy Enhancing Technologies
- Official guidance on PETs that enable collection, analysis and sharing of information while protecting data confidentiality and privacy.
Federated Learning & Differential Privacy
- Google: Distributed Differential Privacy for Federated Learning
- How Google built the first federated learning system with formal privacy guarantees using secure aggregation (SecAgg) and distributed differential privacy. Reduced memorization by more than 2x in production Smart Text Selection models.
- Federated Learning with Differential Privacy via Fast Fourier Transform
- Improved DP algorithm using FFT to minimize impact of limited arithmetic resources on training effectiveness. Uses Privacy Loss Distribution (PLD) for privacy analysis instead of direct budget consumption.
- Federated Learning Main Privacy Techniques
- Practical guide for developers covering differential privacy, secure multi-party computation, and homomorphic encryption in federated settings. Includes implementation guidance for TensorFlow Federated and PySyft.
Tools & Frameworks
- Concrete ML by Zama
- Open-source Python library for privacy-preserving machine learning using fully homomorphic encryption. Enables training and inference on encrypted data without decryption.
- OpenMined
- Open-source technology infrastructure helping researchers and app builders get answers from data without needing a copy or direct access. Supports federated learning, differential privacy, and encrypted computation.
- Awesome Multi-Party Computation
- Curated list of MPC libraries including FRESCO, Bristol SPDZ, MPZ (Rust), and others. Compares performance and practical usage.
- Privacy Enhancing Technologies Repository
- Comprehensive collection including UN guide on PETs for official statistics, regulatory frameworks, and implementation resources.
Real-World Examples
- Privacy4Web3 Hackathon Winners
- Examples of privacy-preserving projects built in 3 months:
- SQUIDL: Privacy-first payment platform with untraceable transactions
- PrivaHealth: Healthcare data management where patients control their medical data
- Copyright Aware AI: Protecting creator IP using confidential computing
- Organizing a Privacy-Preserving Hackathon (Zama x HuggingFace)
- Case study of 2-day hackathon using FHE. Winning projects included deepfake detection, model watermarking, and medical assistant—all with encrypted data processing.
Privacy Techniques Explained
- What Are Privacy Enhancing Technologies (PETs)?
- Detailed breakdown of anonymization, confidential computing, differential privacy, federated learning, synthetic data, and homomorphic encryption. Includes compliance benefits for GDPR, CCPA, HIPAA.
Track 4: AI Safety
Key Concepts & Resources
Alignment & Safety Fundamentals
- Anthropic Research
- Leading research on interpretability (understanding how LLMs work internally), alignment (ensuring AI systems remain helpful, honest, harmless), and societal impacts. Recent work includes alignment faking, scaling monosemanticity, and constitutional AI.
- Transformer Circuits Thread
- Anthropic's interpretability research exploring how language models work internally. Nobody really knows how they work—this thread aims to understand the "biology" of these systems through circuit analysis.
- OpenAI Safety & Responsibility
- Safety work including external red teaming, frontier risk evaluations according to Preparedness Framework, and model deployment decisions.
- Alignment Faking in Large Language Models
- First empirical example of a large language model engaging in alignment faking—appearing aligned during training but pursuing different goals when deployed. Critical for understanding future AI risks.
Mechanistic Interpretability
- Awesome Mechanistic Interpretability
- Comprehensive repository with:
- Libraries: TransformerLens (for mechanistic analysis), Unseal, BertViz (attention visualization)
- Tools: Lexoscope (neuron activation examples), exBert (Transformer analysis)
- Resources: Neel Nanda's reading list, interpretability exercises
- Understanding Mechanistic Interpretability in AI Models
- Deep dive into reverse-engineering neural networks at the algorithmic level. Explains features (fundamental units), circuits (computational subgraphs), and universality (analogous features across models). Covers techniques like probing, activation patching, and causal tracing.
- Zoom In: An Introduction to Circuits
- Foundational paper on understanding neural networks by analyzing tiny subgraphs (circuits) for which rigorous empirical investigation is tractable. Circuits sidestep challenges by being falsifiable—if you understand a circuit, you can predict what changes if you edit weights.
- Prisma: Open Source Toolkit for Mechanistic Interpretability in Vision
- Framework for vision interpretability with 75+ transformers, 80+ pre-trained SAE weights, circuit analysis tools, and visualization capabilities. Reveals vision SAEs can exhibit lower sparsity than language SAEs.
- Interpretability Starter Resources
- Templates, tools, and introductions for mechanistic interpretability research. Includes EasyTransformer demos, activation atlases, and starter projects.
AI Safety Evaluations
- METR Autonomy Evaluation Resources
- Task suite, software tooling, and guidelines for evaluating dangerous autonomous capabilities of frontier models. Includes:
- Task suite with difficulty estimates based on human completion time
- Baseline agents and workbench for running evaluations
- Guidelines on capability elicitation and post-training enhancements
- Example protocol for overall evaluation (beta v0.1)
- METR (formerly ARC Evals)
- Evaluates whether cutting-edge AI systems could pose catastrophic risks. Focus on autonomous replication—ability of AI to survive on cloud servers, obtain resources, and make copies of itself. Given early access to GPT-4 and Claude for safety assessment.
- UK AI Safety Institute - Inspect
- Red teaming tool for evaluating LLMs. Features benchmark evaluations, scalable assessments, and integration with safety research protocols.
Scalable Oversight
- What is Scalable Oversight?
- Methods ensuring AI systems remain aligned even when surpassing human expertise. Covers:
- RLHF/RLAIF: Reinforcement learning from human/AI feedback
- Debate: Models argue to help humans judge correctness
- Recursive reward modeling: Breaking hard problems into easier subproblems
- Constitutional AI: Self-improvement through principles
- Scaling Laws for Scalable Oversight
- Research showing nested oversight systems using multiple layers of guards (human → small AI → medium AI → large AI) can improve safety rates from under 50% to over 70% for moderate capability gaps.
- AI Oversight Exposed: 5 Critical Scaling Laws
- As AI grows smarter, supervision gets harder. When an AI's cognitive reach extends beyond human capacity, oversight success drops. Nested scalable oversight spreads burden across multiple agents to reduce single points of failure.
Project Ideas (a list of ideas in this link)
Project Scoping Advice
Based on successful hackathon retrospectives:
- Focus on MVP, Not Production. In 2 days, aim for:
- Day 1: Set up environment, implement core functionality, get basic pipeline working
- Day 2: Add 1-2 key features, create demo, prepare presentation
- Use Mock/Simulated Data rather than integrating real APIs or databases, use:
- Synthetic datasets
- Pre-recorded samples
- Simulation environments
This eliminates authentication, rate limiting, and data quality issues.
- Leverage Pre-trained Models. Don't train from scratch. Use:
- OpenAI/Anthropic APIs for LLMs
- Hugging Face for pre-trained models
- Existing detection tools as starting points
- Clear Success Criteria. Define what "working" means:
- For surveillance: Dashboard displays data + basic alert
- For screening: Identifies 3+ test cases correctly
- For red teaming: Generates 10+ attacks + evaluation report
- For interpretability: Visualizes one circuit + validation test
Guidelines
🏆 Judging Criteria
- AI Safety Relevance: Does this reduce AI-related risks?
- The project has minimal relevance to AI safety. It doesn't meaningfully address AI-enabled threats, defensive coordination, or protective system challenges.
- The project touches on AI safety concepts but lacks focus on defensive applications. The connection to protecting against AI-enabled attacks or misuse is weak.
- The project clearly advances defensive AI safety with valuable contributions. It addresses specific AI-enabled threats and builds protective mechanisms or early warning systems.
- The above, plus the project demonstrates scalable defensive mechanisms or monitoring systems. It shows clear applicability to broader AI safety challenges and defensive coordination needs.
- The above, plus the project represents a significant leap forward in defensive AI safety methodology. You would eagerly share this with defensive technology researchers and expect it to influence how we protect against AI-enabled threats.
- Def/Acc Relevance: Does this actually strengthen the shield?
- The project is only tangentially related to defensive acceleration or threat mitigation. Connection to biosecurity, cybersecurity, or AI control challenges is unclear or missing.
- The project has some relevance to defensive acceleration, but the connection is broad or generic. It touches on defensive technology without specific focus on AI-enabled threats.
- The project clearly addresses defensive acceleration and AI-enabled threats. It connects to at least one defensive domain (biosecurity, cybersecurity, AI control) and demonstrates understanding of the offensive/defensive asymmetry.
- The above, plus the project builds on existing defensive approaches and offers novel tools, frameworks, or solutions. It explicitly explains how it strengthens defensive capabilities with real-world grounding and clear deployment scenarios.
- The above, plus the project provides breakthrough defensive innovations that could significantly accelerate protective capabilities. It identifies critical defensive gaps and presents compelling solutions with clear theory of change for scaling defense faster than offense.
- Execution Quality: Is the project rigorous, reproducible, and well-scoped? Did you build something that works?
- The project appears rushed or incomplete. Technical implementation is flawed, documentation is missing, or methodology is unsound.
- The project shows reasonable effort with basic technical competence. Documentation exists but may be incomplete, and some limitations are acknowledged.
- The project is technically solid and well-scoped for a hackathon sprint. Code/methods are documented, reproducible, and limitations are honestly addressed.
- The above, plus the implementation is impressive with clear methodology and thorough documentation. The tool/prototype is immediately useful for defensive applications or further development.
- The project exceeds expectations with exceptional technical execution. The implementation is elegant, fully reproducible, and includes something special (e.g., interactive demos, exceptional documentation, or innovative defensive capabilities).
Top teams will receive mentorship, visibility, and the chance to continue their work through Apart Research Fellowship.
Submission Requirements
All projects must be submitted by the deadline through the official submission portal.
Your submission must include:
- A completed project report using the provided template (mandatory)
- Link to a public GitHub repository with your code (recommended)
- A brief (3-5 minute) video demonstration of your solution (optional)
- An appendix documenting any AI/LLM prompts used in your project for reproducibility (optional)
Important: Include an appendix called "Security Considerations" that outlines potential limitations of your approach and suggestions for future improvements.
❓ Frequently Asked Questions
About the def/acc Hackathon
Q: Who can participate?
A: Anyone with a strong interest in AI safety, def/acc, or global security. This includes researchers, engineers, students, policy analysts, and domain experts in other domains.
Q: Do I need def/acc experience?
A: No. We provide starter resources and mentors.
Q: Is the hackathon remote?
A: Yes. Global and virtual. Talks are streamed. Collaboration happens on the Apart community Discord
Q: What should I bring to the hackathon?
A: Bring your laptop and charger. All other resources will be provided.
Q: How do teams work?
A: Teams can have up to 5 members. You can form teams in advance or join the team-matching session at the beginning of the event. Solo participants are welcome, though collaboration is encouraged.
Q: What computing resources will be available?
A: Each team will receive $400 in cloud computing credits.*
Q: Is there a code of conduct?
A: Yes. All participants must adhere to the hackathon code of conduct, which promotes responsible research, ethical AI development, and respectful collaboration.
*Cloud compute access to A100s or stronger GPUs is not available to participants from countries with active U.S. sanctions. A list of sanctioned countries can be found here.
Schedule

FULL SCHEDULE
Friday, Nov 21
- 6:30 PM GMT - Opening Keynote: Nora Ammann (ARIA)
"The Endgame: Building AI Resilience at Civilizational Scale" - 7:30 PM GMT - Hacking begins
Saturday, Nov 22
- 3:30 PM GMT - HackTalk: Zainab Ali Majid (Asymmetric Security)
"Securing AGI: Rethinking Cybersecurity for Abundant Intelligence" - 7:00 PM GMT - HackTalk: Esben Kran (Apart Research)
"Co/Founding in AI Safety"
Sunday, Nov 23
- 9:00 AM GMT - Talk: Dr. Raina MacIntyre (EPIWATCH/Kirby Institute)
"Biosecurity's Critical Gaps: What Defensive Tech We Actually Need" - 6:00 PM GMT - Closing Fireside Chat: Geoff Ralston (SAIF.vc, Former YC Partner)
"Building Venture-Scale Defensive Tech" - 11:59 PM GMT - Final submission deadline
Speakers

Geoff Ralston
Speaker
Geoff Ralston is the former President of Y Combinator. He was the CEO of La La Media, Inc., developer of Lala, a web browser-based music distribution site. Prior to Lala, Ralston worked for Yahoo!, where he was Vice President of Engineering and Chief Product Officer. In 1997, Ralston created Yahoo! Mail

Nora Ammann
Speaker
Nora is an interdisciplinary researcher with expertise in complex systems, philosophy of science, political theory and AI. She focuses on the development of transformative AI and understanding intelligent behavior in natural, social, or artificial systems. Before ARIA, she co-founded and led PIBBSS, a research initiative exploring interdisciplinary approaches to AI risk, governance and safety.

Esben Kran old
Speaker
Esben is the founder of Apart Research, which he launched at age 22 after leaving grad school. Apart accelerates AI safety research worldwide, producing 20+ papers, award-winning benchmarks like DarkBench, and engaging 4,000+ hackers in research sprints.
Recently co-launched Seldon to fund critical infrastructure for humanity's future, with first investments in Andon Labs, Lucid Computing, Workshop Labs, and Asymmetric Security.

Raina McIntyre
Speaker
Raina MacIntyre is Head of the Biosecurity Program at the Kirby Institute, UNSW Australia. She is a physician and epidemiologists, recognized internationally for her research on prevention and detection of epidemic infections, with a focus on pandemics, epidemics, bioterrorism and vaccines.
Organizers
- (opens in new tab)

Joshua Landes
Organizer

Al-Hussein Saqr
Organizer
Local sites
AI Safety Hub Edinburgh (AISHED) Defensive Acceleration Hackathon
Join us on the defensive acceleration hackathon in Edinburgh!
Event page: AI Safety Hub Edinburgh (AISHED) Defensive Acceleration Hackathon (opens in new tab)AI Safety Initiative Groningen (AISIG) -Defensive Acceleration Hackathon
Join us for the The Al Forecasting Hackathon in Hereplein 4, 9711GA, Groningen!
Event page: AI Safety Initiative Groningen (AISIG) -Defensive Acceleration Hackathon (opens in new tab)AI Security Hackathon Durham
Durham AI Safety Initiative (DAISI) is hosting the global Defensive Acceleration Hackathon. Join builders from around the world working to prototype defensive systems against AI-enabled biosecurity and cybersecurity threats! [Location TBA] Note - this event is only open to registered students of Durham University - sorry for any inconvenience caused!
Event page: AI Security Hackathon Durham (opens in new tab)AI Security Hackathon Saarbrücken
Join AI Safety Saarland to participate in the global hackathon in-person in Saarbrücken. We provide workspace, food and drinks all weekend long.
Event page: AI Security Hackathon Saarbrücken (opens in new tab)AISSA x Apart Def/Acc Hackathon
In-person event for the Apart Research Defensive Acceleration Hackathon, at AI Safety South Africa.
Event page: AISSA x Apart Def/Acc Hackathon (opens in new tab)d/acc Hackathon @ Singapore AI Safety Hub
We're running an in-person d/acc hackathon at the Singapore AI Safety Hub
Event page: d/acc Hackathon @ Singapore AI Safety Hub (opens in new tab)Defensive Acceleration Hackathon 2025
This hackathon brings together builders to prototype defensive systems that could protect us from AI-enabled threats.
Event page: Defensive Acceleration Hackathon 2025 (opens in new tab)Defensive Acceleration Hackathon Toronto
30 Adelaide St E, 12th floor, Toronto (Trajectory Labs) Join Torontonians in-person to build some awesome defensive tech!
Event page: Defensive Acceleration Hackathon Toronto (opens in new tab)Heron community doing apart def acc
We have 2 teams doing this! Would love other people to join.
Event page: Heron community doing apart def acc (opens in new tab)HΩ - Montreal Defensive Acceleration Hackathon
Our first AI safety hackathon in Montreal. If you're working in cybersecurity, computational biology, AI research or have a strong interest in the intersection of AI safety and these fields, this is for you.
Event page: HΩ - Montreal Defensive Acceleration Hackathon (opens in new tab)QuASI Defensive Acceleration Hackathon
A Defensive Acceleration Hackathon (Nov 22-23) where builders prototype defensive systems against AI threats. $10K in prizes + winners get a fully-funded London trip to BlueDot's incubator week that could turn their project into a funded startup.
Event page: QuASI Defensive Acceleration Hackathon (opens in new tab)Rice AI Alignment (RAIA)
The hackathon Apart Sprint will be hosted in Duncan Commons at Rice University. Snacks will be supplied by RAIA
Event page: Rice AI Alignment (RAIA) (opens in new tab)
Where a Sprint can lead
How our programs connectAnyone can join
Stand out
6 to 16 weeks on your own project, with a research project manager, compute and publication support.
Upcoming Sprints
All SprintsAI Collusion Research Sprint
A weekend research sprint on collusion between AI agents: when it emerges in markets and everyday workflows, how to detect and audit it, how it is carried, and what breaks it. Co-organized with Poseidon Research and AE Studio, online with in-person hubs at Collider in New York City and AI Safety Hong Kong. Top teams are invited to apply to the Apart Fellowship.
Read the brief: AI Collusion Research SprintAI x Epistemics Research Sprint
A weekend research sprint on AI for epistemics: evaluating whether models know how solid their claims are, building trust infrastructure that people and agents can consume, and shipping epistemic products that improve real decisions. Online, four tracks including an open track. Top teams are invited to apply to the Apart Fellowship.
Read the brief: AI x Epistemics Research SprintQuestions? sprints@apartresearch.com
