News and updates
Research releases, Sprint round-ups, researcher spotlights and the newsletter archive.
- Community
Explaining the Apart Research Fellowships
And introducing our brand new Partnered Fellowships
Read article: Explaining the Apart Research Fellowships - Research
Problem Areas in Physics and AI Safety
We outline five key problem areas in AI safety for the AI Safety x Physics hackathon.
Read article: Problem Areas in Physics and AI Safety - Newsletter
Apart: Two Days Left of our Fundraiser!
Last call to be part of the community that contributed when it truly counted
Read article: Apart: Two Days Left of our Fundraiser! - Newsletter
Apart: Fundraiser Extended!
We've received another $462,276 since our last newsletter, making the total $597k of our $955k goal!
Read article: Apart: Fundraiser Extended! - Newsletter
Apart: Fundraiser Update!
Fundraiser Momentum Builds with Overwhelming Community Support!
Read article: Apart: Fundraiser Update! - Community
Beyond Monolithic AI: The Case for an Expert Orchestration Architecture
Replace Monolithic AIs with an alternative architecture "Expert Orchestration" that intelligently selects from thousands of specialized models based on query requirements, to democratize AI development, increase transparency, and decrease safety risk.
Read article: Beyond Monolithic AI: The Case for an Expert Orchestration Architecture - Newsletter
Apart News: Transformative AI Economics
This week we launched our AI Economics Hackathon as we continue to think about the potentially transformative effects of AI on the global economy.
Read article: Apart News: Transformative AI Economics - Research
Engineering a World Designed for Safe Superintelligence
Esben explains how we go about "Engineering a World Designed for Safe Superintelligence" at Paris' AI Action Summit.
Read article: Engineering a World Designed for Safe Superintelligence - Newsletter
Apart News: Our Biggest Event Ever
This week we have details of our Control Hackathon and a writeup from our biggest event ever.
Read article: Apart News: Our Biggest Event Ever - Community
Women in AI Safety: Hackathon Round-Up
Our Women in AI Safety Hackathon was Apart Research's biggest event ever. Read about the winners here.
Read article: Women in AI Safety: Hackathon Round-Up - Community
Mapping AI Safety Research: An Open-Source Knowledge Graph
A tool to map the sprawling landscape of AI alignment research
Read article: Mapping AI Safety Research: An Open-Source Knowledge Graph - Newsletter
Apart News: San Francisco Edition
This week we have been in San Francisco for our Apart Retreat, where we attended conferences, saw old friends, and visited other AI labs to talk about frontier AI.
Read article: Apart News: San Francisco Edition - Newsletter
Apart News: ICLR Awards & Women in AI Safety
This week, we celebrate ICLR conference oral awards for two of our papers, launch our Women in AI Safety hackathon, and more.
Read article: Apart News: ICLR Awards & Women in AI Safety - Research
Uncovering Model Manipulation with DarkBench
Apart Research developed DarkBench to uncover dark patterns - application design practices that manipulate a user’s behavior against their intention - in some of the world's most popular in LLMs.
Read article: Uncovering Model Manipulation with DarkBench - Research
Studio Progress Report
We are happy to share the significant progress made by the first batch of Apart Research's Studio projects.
Read article: Studio Progress Report - Newsletter
Apart News: Esben at IASEAI & Studio Progress Report
This week Esben gave a talk in Paris and our inaugural Studio Progress Report is released soon.
Read article: Apart News: Esben at IASEAI & Studio Progress Report - Newsletter
Apart News: Paris AI Summit & Catching Hackers
This week some of the team are in Paris & we have just published an Apart Lab Studio research blog about catching AI hackers.
Read article: Apart News: Paris AI Summit & Catching Hackers - Community
AI Safety Entrepreneurship Hackathon Round-Up
In his Hackathon Round-Up we check out the winners of our AI Entrepreneurship Hackathon.
Read article: AI Safety Entrepreneurship Hackathon Round-Up - Newsletter
Apart News: AI Entrepreneurship & New Research
This week we reveal our AI Startup Hackathon winners and have a look at the Apart Lab paper just accepted to ICLR's 2025 conference.
Read article: Apart News: AI Entrepreneurship & New Research - Newsletter
Apart News: Exclusive Interview with Interpretability Insider
Myra reveals how Goodfire's groundbreaking API enabled 200+ researchers at Apart's global hackathon to advance AI interpretability, demonstrating new ways to make AI systems more transparent and controllable.
Read article: Apart News: Exclusive Interview with Interpretability Insider - Research
Behind the Features: Goodfire's Interpretability Tools in Action
Goodfire's Myra reveals how their groundbreaking API enabled 200+ researchers at Apart's global hackathon to advance AI interpretability, demonstrating new ways to make AI systems more transparent and controllable.
Read article: Behind the Features: Goodfire's Interpretability Tools in Action - Research
Promising results from Latent Adversarial Training
Apart Research's newest research achieves promising results from Latent Adversarial Training.
Read article: Promising results from Latent Adversarial Training - Newsletter
Apart News: new LAT research just dropped
In this week's Apart News we look over promising new LAT research and get a Hackathon insider's account from Archana.
Read article: Apart News: new LAT research just dropped - Community
Inside the first AI Policy Hackathon at Johns Hopkins
Johns Hopkins University hosted its first AI Policy Hackathon in partnership with us at Apart Research. Here's what participants and organizers had to say about bridging the gap between technology and policy.
Read article: Inside the first AI Policy Hackathon at Johns Hopkins - Community
Apart in 2024
2024 was the biggest and most impactful year of Apart Research so far.
Read article: Apart in 2024 - Research
AI Hackers in the Wild: LLM Agent Honeypot
This Apart Lab Studio research blog attempts to ascertain the current state of AI-powered hacking in the wild through an innovative 'honeypot' system designed to detect LLM-based attackers.
Read article: AI Hackers in the Wild: LLM Agent Honeypot - Newsletter
Apart News: Hackathons in 2025 PREVIEW
In this week's Apart News we preview some of the Hackathons we are most excited for in 2025.
Read article: Apart News: Hackathons in 2025 PREVIEW - Newsletter
Sparse Autoencoder Hackathon
Our Hackathon round-up showcases our global sprints community.
Read article: Sparse Autoencoder Hackathon - Research
Rethinking CyberSecEval: An LLM-Aided Approach to Evaluation Critique
Apart Research's newest paper looks at LLM-assisted benchmark analysis.
Read article: Rethinking CyberSecEval: An LLM-Aided Approach to Evaluation Critique - Newsletter
Apart News: our research at NeurIPS
In this week's Apart News we are at NeurIPS in Canada.
Read article: Apart News: our research at NeurIPS - Newsletter
Apart News: *NEW VIDEO* Jacob Haimes on working at Apart
In this week's Apart News we have a *brand new* video edition of our Researcher Spotlight series.
Read article: Apart News: *NEW VIDEO* Jacob Haimes on working at Apart - Newsletter
Apart News: 2024 was our biggest year yet
In this week's Apart News we invite you to revisit Apart Research's incredible 2024 with us.
Read article: Apart News: 2024 was our biggest year yet - Newsletter
Apart News: how impactful are we?
In this week's edition of Apart News we ask just how impactful a donation is to Apart Research and take a look at the ability of LLMs to predict neuroscience results.
Read article: Apart News: how impactful are we? - Newsletter
Apart News: NEW Papers, Elections & Goodfire
In this week's edition of Apart News we have two *NEW* papers, the EA Forum's Donation Election, and news from another two of our hackathons this week, with Howard University and Goodfire.
Read article: Apart News: NEW Papers, Elections & Goodfire - Research
Testing LLMs' ability to find security flaws in Cryptographic Protocols
Apart Research's newest paper offers a systematic way to evaluate how well Large Language Models (LLMs) can identify vulnerabilities in cryptographic protocols.
Read article: Testing LLMs' ability to find security flaws in Cryptographic Protocols - Community
How impactful is donating to Apart Research?
Co-Director Esben gives us his thoughts on how impactful donating to Apart Research is.
Read article: How impactful is donating to Apart Research? - Newsletter
Apart News: Announcing Apart Lab Studio
In this week's edition of Apart News, we are super excited to announce our brand new AI Safety program: Apart Lab Studio. And of course, it wouldn't be an edition of Apart News if we didn't have a new paper to share, too.
Read article: Apart News: Announcing Apart Lab Studio - Community
Announcing Apart Lab Studio
Our new Apart Lab Studio is designed to bridge the gap between weekend hackathon projects and a fully-fledged AI Safety research career.
Read article: Announcing Apart Lab Studio - Newsletter
Apart News: Ale, Cash Prizes & the UK’s AISI
This week's edition of Apart News features our newest Researcher Spotlight, informs readers of some of our funding success, and a new cash bounty program from the UK's AI Safety Institute.
Read article: Apart News: Ale, Cash Prizes & the UK’s AISI - Spotlight
Researcher Spotlight: Alexandra Abbas
Alexandra Abbas is currently a Fellow with us here at Apart, where she is focusing on studying the robustness of adversarial fine tuning techniques against the ablation of the refusal feature.
Read article: Researcher Spotlight: Alexandra Abbas - Newsletter
Apart News: Esben, Winning Sprints & ‘3cb’
This week's edition of Apart News has excerpts from Esben's blog, goes through our winning AI Policy sprints, takes a closer look at our new '3cb' benchmark, and more.
Read article: Apart News: Esben, Winning Sprints & ‘3cb’ - Research
Esben on AGI, 'Sentware', and Confident optimism
Esben Kran gives us some of his thoughts on ideas relevant to AI safety, decision-making, and more.
Read article: Esben on AGI, 'Sentware', and Confident optimism - Research
‘3cb’: The Catastrophic Cyber Capabilities Benchmark
Apart Research's newest paper, Catastrophic Cyber Capabilities Benchmark (3CB): Robustly Evaluating LLM Agent Cyber Offense Capabilities, creates a novel cyber offense capability benchmark.
Read article: ‘3cb’: The Catastrophic Cyber Capabilities Benchmark - Newsletter
AI Policy Hackathon in Washington D.C.
Our Hackathon round-up showcases our global 'sprints' community.
Read article: AI Policy Hackathon in Washington D.C. - Newsletter
Apart News: Finn, Cyber Offense & Johns Hopkins
This week's edition of Apart News introduces our new AI Safety Startup project which had its first roundtable with would-be founders, we share our work on cyber offense evaluations for superintelligent AI, and our AI Policy Hackathon is finally kicking off this weekend.
Read article: Apart News: Finn, Cyber Offense & Johns Hopkins - Newsletter
Apart News: Clement, Benchmarks & D.C.
This week's edition of Apart News introduces Clement Neo our Research Assistant who has his profile published in the newest edition of Researcher Spotlight, announces our new paper and accompanying blog on Benchmark Inflation, looks at the Minecraft-themed Agent Security Sprint winner, and gives more information about our new Hackathon in Washington D.C.
Read article: Apart News: Clement, Benchmarks & D.C. - Research
Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts
Apart Research's newest paper finds that many public benchmarks may no longer provide accurate evaluations due to the inclusion of test data in training datasets.
Read article: Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts - Spotlight
Researcher Spotlight: Clement Neo
"Hi! Yes I’m originally from Singapore. I’m currently a Research Assistant at Apart Research, where I focus on technical AI safety and mechanistic interpretability. I grew up with a keen interest in tech and got my first computer when I was just three years old! My dad is an engineer, and that had a big influence on me early on."
Read article: Researcher Spotlight: Clement Neo - Spotlight
Researcher Spotlight: Akash Kundu
This is the first of our Researcher Spotlights: which will shine a light on the global community behind Apart Research. We have so many incredible and talented researchers at Apart Research, and this series will help you to get to know them a little better. First up we have Akash Kundu, Apart Research's Lab Fellow.
Read article: Researcher Spotlight: Akash Kundu - Newsletter
Apart News: Researcher Spotlight, New Team Member & Bangalore
This week's edition of Apart News contains the inaugural edition of our Researcher Spotlights, introduces a new Apart Team member, highlights more of our paper and workshop successes, and reveals a new international Jam Site.
Read article: Apart News: Researcher Spotlight, New Team Member & Bangalore - Research
Esben on agent safety research
Agent safety research is difficult because it involves many different types of entities and wide range of vulnerabilities and failure modes. As a result, it’s hard to develop research that generalizes to all agents. However, we need to give it a shot!
Read article: Esben on agent safety research - Newsletter
Apart News: Agents, Submissions & Spain
This week's scorching edition of Apart News begins with our Agent Security Hackathon kicking off this weekend, our new and exciting Jam Site locations announcement, a whole host more of paper submissions shipped by Apart Lab, and updates on our ongoing Apart Balearic Retreat.
Read article: Apart News: Agents, Submissions & Spain - Newsletter
Apart News: New Research, NeurIPS Papers & Team Offsite
We are proud to announce that Apart’s new co-authored paper, Interpreting Learned Feedback Patterns in Large Language Models, has been accepted to the prestigious NeurIPS. Read the full paper here.It is authored by Luke Marks (Apart Research Fellow), Amir Abdullah (Cynch.ai, Apart Research Fellow), Fazl Barez (Apart Research Advisor) and Clement Neo (Apart Research Assistant), Rauno Arike (Apart), David Krueger (Cambridge_Uni), and Phillip Torr (Department of Engineering Sciences, University of Oxford).
Read article: Apart News: New Research, NeurIPS Papers & Team Offsite - Research
Do models really internalize our preferences?
Apart Research's newest paper (alongside academics from the University of Oxford, Cambridge, and Cynch.ai) looks at whether models actually internalize human preferences or not. But why does this matter?
Read article: Do models really internalize our preferences? - Newsletter
Apart News: o1, Awards & Singapore
This week’s edition of Apart News looks at results from a recent hackathon, analysis about OpenAI's o1, and explains why some of our team are in Singapore.
Read article: Apart News: o1, Awards & Singapore - Newsletter
Apart News: AI Startups, India & Concordia
Our community of researchers is so important to us and Apart News will reflect that. This newsletter will usually include this type of content:
Read article: Apart News: AI Startups, India & Concordia - Community
Can startups be impactful in AI safety?
This post details the top projects from our technical AI safety startups hackathon where researchers and entrepreneurs joined from across the world.
Read article: Can startups be impactful in AI safety? - Community
Where we are on for-profit AI safety
Read about how Big Tech's AI race leaves safety in the dust, non-profits struggle to keep up, and the challenges for-profit AI safety ventures must overcome to leverage resources and make a real impact.
Read article: Where we are on for-profit AI safety - Community
Finding Deception in Language Models
This June, Apart Research and Apollo Research joined forces to host the Deception Detection Hackathon, bringing together students, researchers and engineers from around the world to tackle one of the most pressing challenges in AI safety: Preventing AI from deceiving humans.
Read article: Finding Deception in Language Models - Community
Code Red LLM Evaluations Hackathon Wrap Up (METR and Apart)
Our 128 participants submitted more than 200 project ideas, 100 detailed task specifications, and more than 20 complete implementations! In this post, we also get an exclusive interview with one of the winners.
Read article: Code Red LLM Evaluations Hackathon Wrap Up (METR and Apart) - Community
The ultimate guide to AI safety research hackathons
Research hackathons are an amazing way to dive into a new field, collaborate with passionate people, and create impactful projects in just a short weekend.
Read article: The ultimate guide to AI safety research hackathons - Community
Join us at the AI x Democracy research hackathon
Participate online or in-person on the weekend 3rd to 5th May in an exciting and intense AI safety research hackathon focused on demonstrating and extrapolating risks to democracy from real-life threat models.
Read article: Join us at the AI x Democracy research hackathon - Community
Join the AI Evaluation Tasks Bounty Hackathon with METR
In this collaboration between METR and Apart, you get the chance to contribute directly to model evaluations research.
Read article: Join the AI Evaluation Tasks Bounty Hackathon with METR - Community
How to organize a research hackathon
Organizing a hackathon can bring a unique and exciting energy to people interested in AI safety research! This post summarizes how you can organize a successful hackathon.
Read article: How to organize a research hackathon - Spotlight
Researcher Spotlight: Jacob Haimes
"Hi, my name is Jacob Haimes and I am from Boulder, Colorado, which is right at the foothills of the Rocky Mountains. I grew up here in Boulder and when I was applying for colleges, I only applied to one because I knew I wanted to go to CU Boulder."
Read article: Researcher Spotlight: Jacob Haimes - Community
For-profit AI Safety
AI development attracts more than $67 billion in yearly investments, contrasting sharply with the $250 million allocated to AI safety. This gap suggests there's a large opportunity for AI safety to tap into the commercial market. The big question is how do you close that gap?
Read article: For-profit AI Safety - Community
Taking your next steps after a research hackathon
With the research hackathon, your journey into the world of AI safety is definitely not over! Besides the chance to join the Apart Lab Fellowship, we have collected a bunch of resources here for you to dive even deeper into the field!
Read article: Taking your next steps after a research hackathon - Community
Why organize a research hackathon?
There are many reasons to run a hackathon but some of the main ones are that hackathons are an amazing way to engage the local groups in AI security research and create a sense of community.
Read article: Why organize a research hackathon? - Community
Updated quickstart guide for mechanistic interpretability
Written by Neel Nanda, who previously worked on mech interp under Chris Olah at Anthropic, who is currently a researcher on the DeepMind mechanistic interpretability team.
Read article: Updated quickstart guide for mechanistic interpretability - Research
Results from the Scale Oversight hackathon
Check out the top projects from the "Scale Oversight" hackathon hosted in February 2023: Playing games with LLMs, scaling of prompt specificity, and more.
Read article: Results from the Scale Oversight hackathon - Research
Results from the AI testing hackathon
See the winning projects from the AI testing hackathon held in December 2022: Trojan networks, unsupervised latent knowledge representation, and token loss trajectories to target interpretability methods.
Read article: Results from the AI testing hackathon - Research
Results from the language model hackathon
See winning projects from the language model hackathon hosted November 2022: GPT-3 shows sycophancy, OpenAI's flagging is biased, and truthfulness is sensitive to prompt design.
Read article: Results from the language model hackathon - Research
Results from the interpretability hackathon
Read the winning projects from the interpretability hackathon hosted in November 2022: Automatic interpretability, backup backup name mover heads, and "loud facts" in memory editing.
Read article: Results from the interpretability hackathon