
Nov 11 - 13, 2022Online and in person
The Interpretability Hackathon
This Sprint has ended.
- Sign-ups
- 217
Projects from this Sprint are not published on the site.
48 hours intense research in interpretability, the modern AI neuroscience.
Overview
Hosted by Esben Kran, Apart Research, Neel Nanda
Join this AI safety hackathon to find new perspectives on the "brains" of AI!

We will work with multiple research directions and one of the most relevant is mechanistic interpretability. Here, we work towards a researcher understanding of neural networks and why they do what they do.

We provide you with the best starter templates that you can work from so you can focus on creating interesting research instead of browsing Stack Overflow. You're also very welcome to check out some of the ideas already posted!
How to participate
Create a user on the itch.io (this) website and click participate. We will assume that you are going to participate and ask you to please cancel if you won't be part of the hackathon.
Instructions
You will work on research ideas you generate during the hackathon and you can find more inspiration below.
Submission
Everyone will help rate the submissions together on a set of criteria that we ask everyone to follow. You receive 5 projects that you have to rate during the 4 hours of judging before you can judge specific projects (so we avoid selectivity in the judging).
Resources
Inspiration
We have many ideas available for inspiration on the aisi.ai Interpretability Hackathon ideas list. A lot of interpretability research is available on distill.pub, transformer circuits, and Anthropic's research page.
Introductions to mechanistic interpretability
- Catherine Olsson's on getting starter with mechanistic interpretability research
- A video walkthrough of A Mathematical Framework for Transformer Circuits.
- The Transformer Circuits YouTube series
- Jacob Hilton's deep learning curriculum week on interpretability
- An annotated list of good interpretability papers, along with summaries and takes on what to focus on.
- Christoph Molnar's book about traditional interpretability
- Neel's barebones prerequisites for mechanistic interpretability research
See also the tools available on interpretability:
- Redwood Research's interpretability tools: http://interp-tools.redwoodresearch.org/
- The activation atlas: https://distill.pub/2019/activation-atlas/
- The Tensorflow playground: https://playground.tensorflow.org/
- The Neural Network Playground (train simple neural networks in the browser): https://nnplayground.com/
- Visualize different neural network architectures: http://alexlenail.me/NN-SVG/index.html
Digestible research
- Opinions on Interpretable Machine Learning and 70 Summaries of Recent Papers summarizes a long list of papers that is definitely useful for your interpretability projects.
- Distill publication on visualizing neural network weights
- Andrej Karpathy's "Understanding what convnets learn"
- Looking inside a neural net
- 12 toy language models designed to be easier to interpret, in the style of a Mathematical Framework for Transformer Circuits: 1, 2, 3 and 4 layer models, for each size one is attention-only, one has GeLU activations and one has SoLU activations (an activation designed to make the model's neurons more interpretable - https://transformer-circuits.pub/2022/solu/index.html) (these aren't well documented yet, but are available in EasyTransformer)
- Anthropic Twitter thread going through some language model results
Below is a talk by Esben on a principled introduction to interpretability for safety:
Speakers

Esben Kran
Organizer and Keynote Speaker
Esben is the founder of Apart Research, which he launched at age 22 after leaving grad school. Apart accelerates AI safety research worldwide, producing 20+ papers, award-winning benchmarks like DarkBench, and engaging 4,000+ hackers in research sprints.
Recently co-launched Seldon to fund critical infrastructure for humanity's future, with first investments in Andon Labs, Lucid Computing, Workshop Labs, and Asymmetric Security.
Local sites
Aarhus Interpretability Hackathon
effective-altruism-denmark
Once again, Aarhus University will host the Alignment Jam for interpretability in November.
Event page: Aarhus Interpretability Hackathon (opens in new tab)ENS Interpretability Hackathon
Arranged at the ENS Ulm, this jam site is open for the talented students and faculty.
Event page: ENS Interpretability Hackathon (opens in new tab)Georgia Tech Interpretability hackathon
The AI Safety Initiative at Georgia Tech are hosting a small jam site for graduates and undergraduates at the university.
Israel Interpretability Hackathon
We are a group of 6-10 (mainly) hackers, ai "experts" and neuroscientists
Event page: Israel Interpretability Hackathon (opens in new tab)LEAH Hackathon Site
Imperial College, UCL, King's College, and LSE are jointly hosting the hackathon at the UCL EA offices in Regus, Charlotte Street.
Event page: LEAH Hackathon Site (opens in new tab)Online & Global Hackathon
GatherTown will be open internationally for the duration of the hackathon on GatherTown and we encourage virtual attendees to join there.
Event page: Online & Global Hackathon (opens in new tab)Prague Interpretability hackathon
Join us in Fixed Point in Prague - Vinohrady, Koperníkova 6 for a weekend research sprint in ML interpretability!
Event page: Prague Interpretability hackathon (opens in new tab)Tallinn EA jam site
Estonia EA (Efektiivne Altruism) is hosting the local jam site in Tallinn at the University.
Event page: Tallinn EA jam site (opens in new tab)
Where a Sprint can lead
How our programs connectAnyone can join
Stand out
6 to 16 weeks on your own project, with a research project manager, compute and publication support.
Upcoming Sprints
All SprintsAI Collusion Research Sprint
A weekend research sprint on collusion between AI agents: when it emerges in markets and everyday workflows, how to detect and audit it, how it is carried, and what breaks it. Co-organized with Poseidon Research and AE Studio, online with in-person hubs at Collider in New York City and AI Safety Hong Kong. Top teams are invited to apply to the Apart Fellowship.
Read the brief: AI Collusion Research SprintAI x Epistemics Research Sprint
A weekend research sprint on AI for epistemics: evaluating whether models know how solid their claims are, building trust infrastructure that people and agents can consume, and shipping epistemic products that improve real decisions. Online, four tracks including an open track. Top teams are invited to apply to the Apart Fellowship.
Read the brief: AI x Epistemics Research SprintQuestions? sprints@apartresearch.com
