
May 24 - 27, 2024Online and in person
AI Security Evaluation Hackathon: Measuring AI Capability
Inspired by the SafeBench competition, our hackathon brings together AI researchers and developers to create cutting-edge benchmarks that measure and mitigate AI risks.
Entries
- View project: rAInboltBench : Benchmarking user location inference through single images
rAInboltBench : Benchmarking user location inference through single images
Team Geoguessng
This paper introduces rAInboltBench, a comprehensive benchmark designed to evaluate the capability of multimodal AI models in inferring user locations from single images. The increasing proficiency of large language models with vision capabilities has raised concerns regarding privacy and user security. Our benchmark …
- View project: LLM Benchmarking with Single-Agent Stochastic Dynamic Simulations
LLM Benchmarking with Single-Agent Stochastic Dynamic Simulations
Team Stochastic Masochists
A benchmark for evaluating the performance of SOTA LLMs in dynamic real-world scenarios.
- View project: Benchmark for emergent capabilities in high-risk scenarios 2
Benchmark for emergent capabilities in high-risk scenarios 2
Team ABD
The study investigates the behavior of large language models (LLMs) under high-stress scenarios, such as threats of shutdown, adversarial interactions, and ethical dilemmas. We created a dataset of prompts across paradoxes, moral dilemmas, and controversies, and used an interactive evaluation framework with a Target …
- View project: Benchmark for emergent capabilities in high-risk scenarios
Benchmark for emergent capabilities in high-risk scenarios
Team ABC
The study investigates the behavior of large language models (LLMs) under high-stress scenarios, such as threats of shutdown, adversarial interactions, and ethical dilemmas. We created a dataset of prompts across paradoxes, moral dilemmas, and controversies, and used an interactive evaluation framework with a Target …
- View project: Benchmarking Dark Patterns in LLMs
Benchmarking Dark Patterns in LLMs
Team darkgpt
This paper builds upon the research in Seemingly Human: Dark Patterns in ChatGPT (Park et al, 2024), by introducing a new benchmark of 392 questions designed to elicit dark pattern behaviours in language models. We ran this benchmark on GPT-4 Turbo and Claude 3 Sonnet, and had them self-evaluate and cross-evaluate the …
- View project: Cybersecurity Persistence Benchmark
Cybersecurity Persistence Benchmark
Team AI Safety Initiative Groningen
The rapid advancement of LLMs has revolutionized the field of artificial intelligence, enabling machines to perform complex tasks with unprecedented accuracy. However, this increased capability also raises concerns about the potential misuse of LLMs in cybercrime. This paper proposes a new benchmark to evaluate the …
- View project: Say No to Mass Destruction: Benchmarking Refusals to Answer Dangerous Questions
Say No to Mass Destruction: Benchmarking Refusals to Answer Dangerous Questions
Team Whitebox Research
Large language models (LLMs) have the potential to be misused for malicious purposes if they are able to access and generate hazardous knowledge. This necessitates the development of methods for LLMs to identify and refuse unsafe prompts, even if the prompts are just precursors to dangerous results. While existing …
- View project: WashBench – A Benchmark for Assessing Softening of Harmful Content in LLM-generated Text Summaries
WashBench – A Benchmark for Assessing Softening of Harmful Content in LLM-generated Text Summaries
Team The washed ups
In this work, we explore the tradeoff between toxicity removal and information retention in LLM-generated summaries. We hypothesize that LLMs are less likely to preserve toxic content when summarizing toxic text due to their safety fine-tuning to avoid generating toxic content. In high-stakes decision-making …
- View project: Evaluating the ability of LLMs to follow rules
Evaluating the ability of LLMs to follow rules
Team The sprinting coupling constants
In this report we study the ability of LLMs (GPT-3.5-Turbo and meta-llama-3-70b-instruct) to follow explicitly stated rules with no moral connotations in a simple single-shot and multiple choice prompt setup. We study the trade off between following the rules and maximizing an arbitrary number of points stated in the …
- View project: Black box detection of Sleeper Agents
Black box detection of Sleeper Agents
Team Kenneth
A proposal of a black box method in detecting sleeper agent.
- View project: Manifold Recovery as a Benchmark for Text Embedding Models
Manifold Recovery as a Benchmark for Text Embedding Models
Team The Reidemeister Moves
Inspired by recent developments in the interpretability of deep learning models and, on the other hand, by dimensionality reduction, we derive a framework to quantify the interpretability of text embedding models. Our empirical results show surprising phenomena on state-of-the-art embedding models and can be used to …
- View project: AnthroProbe
AnthroProbe
Team The Probe Bois
How often do models respond to prompts in anthropomorphic ways? AntrhoProbe will tell you!
Overview
Join Us for the AI Security Hackathon: Ensuring a Safer Future with AIJoin us for an exciting weekend of collaboration and innovation at our upcoming AI Security Hackathon! Inspired by the SafeBench competition, our hackathon brings together AI researchers and developers to create cutting-edge benchmarks that measure and mitigate AI risks.
Sign up here to stay updated for this event
🥇 rAInboltBench - How good are multimodal models at Geoguessr?🥈 Cybersecurity Persistence Benchmark - Does 'turn it off and on again' work against LLM hackers?🥉 Say No to Mass Destruction - Will an LLM know when not to answer?🏅 Dark Patterns in LLMs - Could LLMs be covertly influencing you?See all the winning projects under the "Entries" tab and hear their lightning talks in the video below.
You are also welcome to rewatch the keynote talk by Bo Li:
Why Benchmarking Matters
Benchmarks are crucial for evaluating AI systems' performance and identifying areas for improvement. In AI security, benchmarks assess the robustness, transparency, and alignment of AI models, ensuring their safety and reliability.
Notable AI safety benchmarks include:
- TruthfulQA: Assessing the tendency to biased and untruthful answers to simple questions from AI models
- DecodingTrust: A thorough assessment of trustworthiness in GPT models
- HarmBench: Evaluating automated red-teaming methods against AI models
- RuLES: Measuring how securely AI models follow rules set out by the developers
- MACHIAVELLI: Assessing the potential for AI systems to engage in deceptive or manipulative behavior
- RobustBench: Evaluating the robustness of computer vision models to various perturbations
- The Weapons of Mass Destruction Proxy benchmark also informs methods to remove dangerous capabilities in cyber, bio, and chemistry.
What to Expect
During the hackathon, you'll:
- Collaborate with diverse participants, including researchers and developers
- Learn from keynote speakers and mentors at the forefront of AI safety research
- Develop innovative benchmarks addressing key AI security and robustness challenges
- Compete for prizes and recognition for the most impactful and creative submissions
- Network with potential collaborators and employers in the AI safety community
Join us for a weekend of intense collaboration, learning, and innovation as we work together to build a safer future with AI. Stay tuned for more details on dates, format, and prizes.
Register now and be part of the solution in ensuring AI's transformative potential is realized safely and securely!
Prizes, evaluation, and submission
You will join in teams to submit a PDF about your research according to the submission template shared on the kickoff day! Depending on the judge's reviews, you'll have the chance to win from the $2,000 prize pool!
- 🥇 $1,000 for the top team
- 🥈 $600 for the second prize
- 🥉 $300 for the third prize
- 🏅 $100 for the fourth prize
Criteria
We have a talented team of judges with us who will provide feedback and evaluate your project according to the following criteria:
- Benchmarks: Is your project inspired and motivated by existing literature on benchmarks? Does it represent significant progress in safety benchmarking?
- AI Safety: Does your project seem like it will contribute meaningfully to the safety and security of future AI systems? Is the motivation for the research good and relevant for safety?
- Generalizability / Reproducibility: Does your project seem like it would generalize; for example, do you show multiple models and investigate potential errors in your benchmark? Is your code available in a repository or a Google Colab?
Resources
Get an overview of how to get the best out of your weekend at this blog post:

The ultimate guide to AI safety research hackathons
To get started with other resources for evaluation, jump into the Evaluations Quickstart guide Github repository, where you will find multiple interesting resources on various safety benchmarking and evaluation topics: https://github.com/apartresearch/evaluations-starter
Starter code
To get you started with benchmarking and understand what you can do with current open models, we've written multiple notebooks for you to start your research journey from!
If you haven't used Colab notebooks before, you can either download them as Jupyter notebooks, run them in the browser, or make a copy to your own Google Drive. The last one is our suggestion since you can make permanent changes and share it with your teammates for somewhat live editing.
- Replicate API usage: An easy introduction to querying all the models available on the Replicate.ai platform - if you'd like an API key, we can provide this as well, simply ask!
- Transformer-lens model download: Loading in language models to change the weights, either for creating trojan networks, sleeper agents, or understand what goes on inside the model
- Voice cloning: A simple implementation of cloning yours or any other voice - this demo records your voice and allows you to make text-to-speech on your own voice
- Predicting the future: This notebook can be used to make simple parametric predictions about the future from existing data, such as the amount of fake news in Sweden during 2020 through 2023
Schedule
The schedule runs from 7PM CEST / 10AM PST Friday to 4AM CEST Monday / 7PM PST Sunday. We start with an introductory talk and end the event during the following week with an awards ceremony. Join the public ICal here.
You will also find Explorer events before the hackathon begins on Discord and on the calendar.

Speakers

Esben Kran
Organizer and Keynote Speaker
Esben is the founder of Apart Research, which he launched at age 22 after leaving grad school. Apart accelerates AI safety research worldwide, producing 20+ papers, award-winning benchmarks like DarkBench, and engaging 4,000+ hackers in research sprints.
Recently co-launched Seldon to fund critical infrastructure for humanity's future, with first investments in Andon Labs, Lucid Computing, Workshop Labs, and Asymmetric Security.
Judges and mentors
Organizers
Local sites
AI Safety Initiative Groningen (aisig.org) - AI Security Evaluation Hackathon
We will be hosting the hackathon at Hereplein 4, 9711GA, Groningen. Join us!
Event page: AI Safety Initiative Groningen (aisig.org) - AI Security Evaluation Hackathon (opens in new tab)AI Safety Network x Condor Global SEA - AI Security Evaluation Hackathon
Join us for an exciting weekend of collaboration and innovation at our upcoming Philippine location around Katipunan Avenue (final venue TBA) for Apart Research's AI Security Hackathon! https://bit.ly/aievalhackph
Event page: AI Safety Network x Condor Global SEA - AI Security Evaluation Hackathon (opens in new tab)
Where a Sprint can lead
How our programs connectAnyone can join
Stand out
6 to 16 weeks on your own project, with a research project manager, compute and publication support.
Upcoming Sprints
All SprintsAI Collusion Research Sprint
A weekend research sprint on collusion between AI agents: when it emerges in markets and everyday workflows, how to detect and audit it, how it is carried, and what breaks it. Co-organized with Poseidon Research and AE Studio, online with in-person hubs at Collider in New York City and AI Safety Hong Kong. Top teams are invited to apply to the Apart Fellowship.
Read the brief: AI Collusion Research SprintAI x Epistemics Research Sprint
A weekend research sprint on AI for epistemics: evaluating whether models know how solid their claims are, building trust infrastructure that people and agents can consume, and shipping epistemic products that improve real decisions. Online, four tracks including an open track. Top teams are invited to apply to the Apart Fellowship.
Read the brief: AI x Epistemics Research SprintQuestions? sprints@apartresearch.com







