Skip to content
Reprogramming AI Models Hackathon

Nov 22 - 25, 2024Online and in person

Reprogramming AI Models Hackathon

This Sprint has ended.

Sign-ups
278
Projects submitted
28
Browse the 28 projects

Sign up for this Sprint

Type N/A if you don’t have one.

Type N/A if you don’t have one.

What you work on, and whether you are open to new roles.

What about this event made you want to take part?

By signing up you agree to our Privacy Policy.

The submission window has closed. Your draft is still here so you can copy it, but it can no longer be submitted.

Submit your project

Project details

A short abstract: what you did, what you found.

PDF, up to 25 MB.

Are you interested in publishing this project? *
Tracks

If this Sprint has numbered tracks, choose the ones your project fits.

PDF, PowerPoint, Keynote or ODP, up to 25 MB.

PNG, JPEG, WebP or GIF, up to 25 MB.

Team details

Team member 1

Leave blank if you don’t have one.

By submitting you agree to the prize terms and our Privacy Policy.

The submission window has closed. Your draft is still here so you can copy it, but it can no longer be submitted.

See upcoming Sprints

Whether you're an AI researcher, a curious developer, or passionate about making AI systems more transparent and controllable, this hackathon is for you. As a participant, you will: Collaborate with experts to create novel AI observability tools Learn about mechanistic interpretability from industry leaders Contribute to solving real-world challenges in AI safety and reliability Compete for prizes and the opportunity to influence the future of AI development Register now and be part of the movement towards more transparent, reliable, and beneficial AI systems. We provide access to Goodfire's SDK/API and research preview playground, enabling participation regardless of prior experience with AI observability.

Entries

Overview

Why This Matters

As AI models become more powerful and widespread, understanding their internal mechanisms isn't just academic curiosity—it's crucial for building reliable, controllable AI systems. Mechanistic interpretability gives us the tools to peek inside these "black boxes" and understand how they actually work, neuron by neuron and feature by feature.

What You'll Get

  • Exclusive Access: Use Goodfire's API to access an interpretable 8B or 70B model with efficient inference.
  • Cutting-Edge Tools: Experience Goodfire's SDK/API for feature steering and manipulation
  • Advanced Capabilities: Work with conditional feature interventions and sophisticated development flows
  • Free Resources: Compute credits for every team to ensure you can pursue ambitious projects
  • Expert Guidance: Direct mentorship from industry leaders throughout the weekend

Project Tracks

1. Feature Investigation

  • Map and analyze feature phenomenology in large language models
  • Discover and validate useful feature interventions
  • Research the relationship between feature weights and intervention success
  • Develop metrics for intervention quality assessment

2. Tooling Development

  • Build tools for automated feature discovery
  • Create testing frameworks for intervention reliability
  • Develop integration tools for existing ML frameworks
  • Improve auto-interpretation techniques

3. Visualization & Interface

  • Design intuitive visualizations for feature maps
  • Create interactive tools for exploring model internals
  • Build dashboards for monitoring intervention effects
  • Develop user interfaces for feature manipulation

4. Novel Research

  • Investigate improvements to auto-interpretation
  • Study feature interaction patterns
  • Research intervention transfer between models
  • Explore new approaches to model steering

Why Goodfire's Tools?

While participants are welcome to use their existing setups, Goodfire's API brings exceptional value to this hackathon as a primary option for participants.

Goodfire provides:

  • Access to a 70B parameter model via API (with efficient inference)
  • Feature steering capabilities made simple through the SDK/API
  • Advanced development workflows including conditional feature interventions

The hackathon serves as a unique opportunity for Goodfire to gather valuable feedback from the developer community on their API/SDK. To ensure all participants can pursue ambitious research projects without constraints, Goodfire is providing free compute credits to every team.

Previous Participant Experiences

"I learned so much about AI Safety and Computational Mechanics. It is a field I have never heard of, and it combines two of my interests - AI and Physics. Through the hackathons, I gained valuable connections and learned a lot from researchers with extensive experience." - Doroteya Stoyanova, Computer Vision Intern

Resources

To ensure you're well-equipped for the Reprogramming AI Models Hackathon, we've compiled a set of resources to support your participation:

  1. Goodfire's SDK/API with hosted inference: Your primary toolkit for the hackathon. Familiarize yourself with our framework for understanding and modifying AI model behavior.
  2. Research Preview Playground
  3. Check out the Jupyter Notebook Quickstart: . In this quickstart, you'll learn how to:
    • Sample from a language model (in this case, Llama 3 8B)
    • Search for exciting features and intervene in them to steer the model
    • Find features by contrastive search
    • Save and load Llama models with steering applied
  4. Tutorial: Visualizing AI Model Internals: Watch this video to understand how to use Goodfire's tools to map and visualize AI model behavior.
  1. The Cognitive Revolution Podcast - Episode on Interpretability. n this episode of The Cognitive Revolution, we delve into the science of understanding AI models' inner workings, recent breakthroughs, and the potential impact on AI safety and control
  2. Auto-interp Paper: This paper applies automation to the problem of scaling an interpretability technique to all the neurons in a large language model.
  3. ARENA Interpretability with SAEs
  4. Gemma Scope: a comprehensive, open suite of sparse autoencoders for language model interpretability.
  5. Neuronpedia: Platform for accelerating research into Sparse Autoencoders
  6. The Geometry of Concepts: Sparse Autoencoder Feature Structure Paper. This paper investigates the structured organization of concept representations within large language models using sparse autoencoders, revealing a multi-scale structure with refined atomic parallelogram forms, modular brain-like spatial features, and anisotropic galaxy-scale distributions with unique eigenvalue properties.
  7. Open Source Replication of Anthropic’s Crosscoder paper for model-diffing
  8. Lesswrong search for SAE

Schedule

Here is the schedule for the Hackathon:
We start with an introductory talk and end the event during the following week with an awards ceremony. Join the public ICal here. You will also find Explorer events, such as collaborative brainstorming and team match-making before the hackathon begins on Discord and in the calendar.

Speakers

  • Neel Nanda

    Neel Nanda

    Speaker & Judge

    Team lead for the mechanistic interpretability team at Google Deepmind and a prolific advocate for open source interpretability research.

  • Joseph Bloom

    Joseph Bloom

    Speaker

    Joseph co-founded Decode Research, a non-profit organization aiming to accelerate progress in AI safety research infrastructure, and is a mechanistic interpretability researcher.

  • Callum McDougall

    Callum McDougall

    Speaker

    ARENA Director and SERI-MATS alumnus specializing in mechanistic interpretability and AI alignment education

  • Esben Kran

    Esben Kran

    Organizer and Keynote Speaker

    Esben is the founder of Apart Research, which he launched at age 22 after leaving grad school. Apart accelerates AI safety research worldwide, producing 20+ papers, award-winning benchmarks like DarkBench, and engaging 4,000+ hackers in research sprints.

    Recently co-launched Seldon to fund critical infrastructure for humanity's future, with first investments in Andon Labs, Lucid Computing, Workshop Labs, and Asymmetric Security.

Judges and mentors

Organizers

Local sites

  • AISIG - Decode AI's Black Box & Engineer Model Behavior Hackathon

    Join us for the Hackathon in Hereplein 4, 9711GA, Groningen!

    Event page: AISIG - Decode AI's Black Box & Engineer Model Behavior Hackathon (opens in new tab)
  • Reprogramming AI Models Hackathon

    This is a collaboration between Warwick AI and Warwick Effective Altruism. We will be hosting groups that wish to participate in the hackathon for the weekend.

    Event page: Reprogramming AI Models Hackathon (opens in new tab)
  • Reprogramming AI Models Hackathon: CAISH

    Cambridge hub for hosting the Reprogramming AI hackathon. Office available with monitors and snacks!

  • Reprogramming AI Models Hackathon: Edinburgh

  • Reprogramming AI Models Hackathon: EPFL hub

    A local hub for the hackathon on the EPFL campus (luma coming soon). We will provide a room, snacks and drinks for the participants.

Where a Sprint can lead

How our programs connect
  1. Sprint

    Anyone can join

    Stand out

  2. Apart Fellowship

    6 to 16 weeks on your own project, with a research project manager, compute and publication support.