Skip to content
The Interpretability Hackathon

Nov 11 - 13, 2022Online and in person

The Interpretability Hackathon

This Sprint has ended.

Sign-ups
217

Projects from this Sprint are not published on the site.

Sign up for this Sprint

Type N/A if you don’t have one.

Type N/A if you don’t have one.

What you work on, and whether you are open to new roles.

What about this event made you want to take part?

By signing up you agree to our Privacy Policy.

The submission window has closed. Your draft is still here so you can copy it, but it can no longer be submitted.

Submit your project

Project details

A short abstract: what you did, what you found.

PDF, up to 25 MB.

Are you interested in publishing this project? *
Tracks

If this Sprint has numbered tracks, choose the ones your project fits.

PDF, PowerPoint, Keynote or ODP, up to 25 MB.

PNG, JPEG, WebP or GIF, up to 25 MB.

Team details

Team member 1

Leave blank if you don’t have one.

By submitting you agree to the prize terms and our Privacy Policy.

The submission window has closed. Your draft is still here so you can copy it, but it can no longer be submitted.

See upcoming Sprints

48 hours intense research in interpretability, the modern AI neuroscience.

Overview

Hosted by Esben Kran, Apart Research, Neel Nanda

Join this AI safety hackathon to find new perspectives on the "brains" of AI!

We will work with multiple research directions and one of the most relevant is mechanistic interpretability. Here, we work towards a researcher understanding of neural networks and why they do what they do.

We provide you with the best starter templates that you can work from so you can focus on creating interesting research instead of browsing Stack Overflow. You're also very welcome to check out some of the ideas already posted!

How to participate

Create a user on the itch.io (this) website and click participate. We will assume that you are going to participate and ask you to please cancel if you won't be part of the hackathon.

Instructions

You will work on research ideas you generate during the hackathon and you can find more inspiration below.

Submission

Everyone will help rate the submissions together on a set of criteria that we ask everyone to follow. You receive 5 projects that you have to rate during the 4 hours of judging before you can judge specific projects (so we avoid selectivity in the judging).

Resources

Inspiration

We have many ideas available for inspiration on the aisi.ai Interpretability Hackathon ideas list. A lot of interpretability research is available on distill.pub, transformer circuits, and Anthropic's research page.

Introductions to mechanistic interpretability

See also the tools available on interpretability:

Digestible research

Below is a talk by Esben on a principled introduction to interpretability for safety:

Speakers

  • Esben Kran

    Esben Kran

    Organizer and Keynote Speaker

    Esben is the founder of Apart Research, which he launched at age 22 after leaving grad school. Apart accelerates AI safety research worldwide, producing 20+ papers, award-winning benchmarks like DarkBench, and engaging 4,000+ hackers in research sprints.

    Recently co-launched Seldon to fund critical infrastructure for humanity's future, with first investments in Andon Labs, Lucid Computing, Workshop Labs, and Asymmetric Security.

  • Neel Nanda

    Neel Nanda

    Speaker & Judge

    Team lead for the mechanistic interpretability team at Google Deepmind and a prolific advocate for open source interpretability research.

Local sites

Where a Sprint can lead

How our programs connect
  1. Sprint

    Anyone can join

    Stand out

  2. Apart Fellowship

    6 to 16 weeks on your own project, with a research project manager, compute and publication support.