
Nov 22 - 25, 2024Online and in person
Reprogramming AI Models Hackathon
Whether you're an AI researcher, a curious developer, or passionate about making AI systems more transparent and controllable, this hackathon is for you. As a participant, you will: Collaborate with experts to create novel AI observability tools Learn about mechanistic interpretability from industry leaders Contribute to solving real-world challenges in AI safety and reliability Compete for prizes and the opportunity to influence the future of AI development Register now and be part of the movement towards more transparent, reliable, and beneficial AI systems. We provide access to Goodfire's SDK/API and research preview playground, enabling participation regardless of prior experience with AI observability.
Entries
- 1st place by peer reviewView project: AutoSteer: Weight-Preserving Reinforcement Learning for Interpretable Model Control
AutoSteer: Weight-Preserving Reinforcement Learning for Interpretable Model Control
Team AI Safety Initiative Groningen
Traditional fine-tuning methods for language models, while effective, often disrupt internal model features that could provide valuable insights into model behavior. We present a novel approach combining Reinforcement Learning (RL) with Activation Steering to modify model behavior while preserving interpretable …
- View project: Classification on Latent Feature Activation for Detecting Adversarial Prompt Vulnerabilities
Classification on Latent Feature Activation for Detecting Adversarial Prompt Vulnerabilities
Team Feature Disruption Lab
We present a method leveraging Sparse Autoencoder (SAE)-derived feature activations to identify and mitigate adversarial prompt hijacking in large language models (LLMs). By training a logistic regression classifier on SAE-derived features, we accurately classify diverse adversarial prompts and distinguish between …
- View project: Utilitarian Decision-Making in Models - Evaluation and Steering
Utilitarian Decision-Making in Models - Evaluation and Steering
Team Byte To Brain
We design an eval based on the Oxford Utilitarianism Scale (OUS) that measures the model’s deontological vs utilitarian preference in a nine question, two factor model. Using this scale, we measure how feature steering can alter this preference, and the ability of the Llama 70B model to impersonate other people’s …
- View project: Steering Swiftly to Safety with Sparse Autoencoders
Steering Swiftly to Safety with Sparse Autoencoders
Team Explaining_Polysemantic_Feature_Learning
We explore using SAEs for unlearning dangerous capabilities in a cheaper and more interpretable way.
- View project: Faithful or Factual? Tuning Mistake Acknowledgment in LLMs
Faithful or Factual? Tuning Mistake Acknowledgment in LLMs
Team Faithfulness Explorers
Understanding the reasoning processes of large language models (LLMs) is crucial for AI transparency and control. While chain-of-thought (CoT) reasoning offers a naturally interpretable format, models may not always be faithful to the reasoning they present. In this paper, we extend previous work investigating chain …
- View project: Feature based unlearning
Feature based unlearning
Team Model memory analysis
An exploration of using features to perform unlearning on answering trivia questions.
- View project: Improving Llama-3-8B-Instruct Hallucination Robustness in Medical Q&A Using Feature Steering
Improving Llama-3-8B-Instruct Hallucination Robustness in Medical Q&A Using Feature Steering
Team Gradients Anatomy
This paper addresses the risks of hallucinations in LLMs within critical domains like medicine. It proposes methods to (a) reduce hallucination probability in responses, (b) inform users of hallucination risks and model accuracy for specific queries, and (c) display hallucination risk through a user interface. Steered …
- View project: Encouraging Chain-of-Thought Reasoning
Encouraging Chain-of-Thought Reasoning
Team BlueDot Impact - Shreyans
Encouraging Chain-of-Thought Reasoning via Feature Steering in Large Language Models
- View project: Investigating Feature Effects on Manipulation Susceptibility
Investigating Feature Effects on Manipulation Susceptibility
Team WAIST 3
In our project, we consider the effectiveness of the AI’s prompt injection protection, and in partic- ular the features that are responsible for providing the bulk of this protection. We prove that the features we identify are responsible for this protection by creating variants of the base model which perform …
- View project: Assessing Language Model Cybersecurity Capabilities with Feature Steering
Assessing Language Model Cybersecurity Capabilities with Feature Steering
Team WAIST1
Searched for the most highly activated weights on cybersecurity questions. Then adjusted these weights to see if the impact multiple choice question answering performance.
- View project: Can we steer a model’s behavior with just one prompt? investigating SAE-driven auto-steering
Can we steer a model’s behavior with just one prompt? investigating SAE-driven auto-steering
Team SAEcret sauce squad
This paper investigates whether Sparse Autoencoders (SAEs) can be leveraged to steer the behavior of models without using manual intervention. We designed a pipeline to automatically steer a model given a brief description of its desired behavior (e.g.: “Behave like a dog”). The pipeline is as follows: 1. We …
- View project: Feature Tuning versus Prompting for Ambiguous Questions
Feature Tuning versus Prompting for Ambiguous Questions
Team LiU AI Safety
This study explores feature tuning as a method to improve alignment of large language models (LLMs). We focus on addressing human psychological fallacies reinforced during the LLM training pipeline. Using sparse autoencoders (SAEs) and the Goodfire SDK, we identify and manipulate features in Llama-3.1-70B tied to …
- View project: Auto Prompt Injection
Auto Prompt Injection
Team WAIST 2
Prompt injection attacks exploit vulnerabilities in how large language models (LLMs) process inputs, enabling malicious behaviour or unauthorized information disclosure. This project investigates the potential for seemingly benign prompt injections to reliably prime models for undesirable behaviours, leveraging …
- View project: BBLLM
BBLLM
Team BBLLM team
This project focuses on enhancing feature interpretability in large language models (LLMs) by visualizing relationships between latent features. Using an interactive graph-based representation, the tool connects co-activated features for specific prompts, enabling intuitive exploration of feature clusters. Deployed as …
- View project: Bias Mitigation
Bias Mitigation
Team Spark
Large Language Models (LLMs) have revolutionized natural language processing, but their deployment has been hindered by biases that reflect societal stereotypes embedded in their training data. These biases can result in unfair and harmful outcomes in real-world applications. In this work, we explore a novel approach …
- View project: Unveiling Latent Beliefs Using Sparse Autoencoders
Unveiling Latent Beliefs Using Sparse Autoencoders
Team Serbot
Language models (LMs) often generate outputs that are linguistically plausible yet factually incorrect, raising questions about their internal representations of truth and belief. This paper explores the use of sparse autoencoders (SAEs) to identify and manipulate features that encode the model’s confidence or belief …
- View project: Clear Thought and Clear Speech: Reducing Grammatical Scope Ambiguity
Clear Thought and Clear Speech: Reducing Grammatical Scope Ambiguity
Team Girzu
With language models starting to be used in fields such as law, unambiguity in wording is an important desideratum in model outputs. I therefore try to find features in Llama-3.1-70B-Instruct that correspond to grammatical scope ambiguity using Goodfire's contrastive feature search tool, and try to steer the model …
- View project: Analyzing Dataset Bias with SAEs
Analyzing Dataset Bias with SAEs
Team Big gamers
We use SAEs to study biases in datasets.
- View project: Let LLM Agents Perform LLM Surgery
Let LLM Agents Perform LLM Surgery
Team Agentic Kittens
This project aimed to create and utilize LLM agents that could perform various mechanistic interventions on other LLMs. A few experiments were conducted ranging from an agent unsteering a mechanistically steered model to a neutral state, to an agent performing mechanistic edits to create a custom LLM as per user …
- View project: Math Speaks All Languages: Enhancing LLM Problem-Solving Across Multilingual Contexts
Math Speaks All Languages: Enhancing LLM Problem-Solving Across Multilingual Contexts
Team Round Tensor
Large language models (LLMs) have shown significant adaptability in tackling various human issues; however, their efficacy in resolving mathematical problems remains inadequate. Recent research has identified steering vectors — hidden attributes that can guide the actions and outputs of LLMs. Nonetheless, the …
- View project: Tentative proposal for AI control with weak supervisors trough Mechanistic Inspection
Tentative proposal for AI control with weak supervisors trough Mechanistic Inspection
Team WTS Team
The project proposes using weak but trusted AI models to supervise powerful, untrusted models by analyzing their internal states via Sparse Autoencoder features. This approach aims to enhance oversight by detecting complex behaviors like deception within the stronger models. Key challenges include managing large-scale …
- View project: Investigate arithmetic features in Multi-lingual LLMs
Investigate arithmetic features in Multi-lingual LLMs
Team one_dos_tres
We investigate the arithmetic related feature activations in Llama3.1 70b model across its 8 supported languages. We use arithmetic-activation strength to compare the 8 languages and unsurprisingly English has the highest strength and Hindi, Thai score the least.
- View project: Explaining Latents in Turing-LLM-1.0-254M with Pre-Defined Function Types
Explaining Latents in Turing-LLM-1.0-254M with Pre-Defined Function Types
Team Latents United
We introduce a novel framework for explaining latents in the Turing-LLM-1.0-254M model based on a predefined set of function types, allowing for a more human-readable “source code” of the model’s internal mechanisms. By categorising latents using multiple function types, we move towards mechanistic interpretability …
- View project: SAGE: Safe, Adaptive Generation Engine for Long Form Document Generation in Collaborative, High Stakes Domains
SAGE: Safe, Adaptive Generation Engine for Long Form Document Generation in Collaborative, High Stakes Domains
Team Benki
Long-form document generation for high-stakes financial services—such as deal memos, IPO prospectuses, and compliance filings—requires synthesizing data-driven accuracy, strategic narrative, and collaborative feedback from diverse stakeholders. While large language models (LLMs) excel at short-form content, generating …
- View project: Recovering Goodfire's SAE feature vectors from their API
Recovering Goodfire's SAE feature vectors from their API
Team Lovkush
In this project, we carry out an early trial to see whether Goodfire’s SAE feature vectors can be recovered using the information available from their API. The strategy tried is: pick a feature of interest, construct a contrastive dataset using Goodfire’s API, then use TransformerLens to get a steering vector for the …
- View project: Edufire - Personalized Education Platform Using LLM Steering
Edufire - Personalized Education Platform Using LLM Steering
Team ThoughtCompass
EduFire is a personalized education platform designed to tailor educational content and assessments to individual user preferences by leveraging the Goodfire API for AI model steering. The platform aims to enhance learner engagement and efficacy by customizing the learning experience according to user-selected …
- View project: Sparse Autoencoders and Gemma 2-2B: Pioneering Demographic-Sensitive Language Modeling for Opinion QA
Sparse Autoencoders and Gemma 2-2B: Pioneering Demographic-Sensitive Language Modeling for Opinion QA
Team momo
This project investigates the integration of Sparse Autoencoders (SAEs) with the gemma 2-2b lan- guage model to address challenges in opinion-based question answering (QA). Existing language models often produce answers reflecting narrow viewpoints, aligning disproportionately with specific demographics. By leveraging …
- View project: Finding Circular Features in Gemma 2 2B
Finding Circular Features in Gemma 2 2B
Overview
Why This Matters
As AI models become more powerful and widespread, understanding their internal mechanisms isn't just academic curiosity—it's crucial for building reliable, controllable AI systems. Mechanistic interpretability gives us the tools to peek inside these "black boxes" and understand how they actually work, neuron by neuron and feature by feature.
What You'll Get
- Exclusive Access: Use Goodfire's API to access an interpretable 8B or 70B model with efficient inference.
- Cutting-Edge Tools: Experience Goodfire's SDK/API for feature steering and manipulation
- Advanced Capabilities: Work with conditional feature interventions and sophisticated development flows
- Free Resources: Compute credits for every team to ensure you can pursue ambitious projects
- Expert Guidance: Direct mentorship from industry leaders throughout the weekend
Project Tracks
1. Feature Investigation
- Map and analyze feature phenomenology in large language models
- Discover and validate useful feature interventions
- Research the relationship between feature weights and intervention success
- Develop metrics for intervention quality assessment
2. Tooling Development
- Build tools for automated feature discovery
- Create testing frameworks for intervention reliability
- Develop integration tools for existing ML frameworks
- Improve auto-interpretation techniques
3. Visualization & Interface
- Design intuitive visualizations for feature maps
- Create interactive tools for exploring model internals
- Build dashboards for monitoring intervention effects
- Develop user interfaces for feature manipulation
4. Novel Research
- Investigate improvements to auto-interpretation
- Study feature interaction patterns
- Research intervention transfer between models
- Explore new approaches to model steering
Why Goodfire's Tools?
While participants are welcome to use their existing setups, Goodfire's API brings exceptional value to this hackathon as a primary option for participants.
Goodfire provides:
- Access to a 70B parameter model via API (with efficient inference)
- Feature steering capabilities made simple through the SDK/API
- Advanced development workflows including conditional feature interventions
The hackathon serves as a unique opportunity for Goodfire to gather valuable feedback from the developer community on their API/SDK. To ensure all participants can pursue ambitious research projects without constraints, Goodfire is providing free compute credits to every team.
Previous Participant Experiences
"I learned so much about AI Safety and Computational Mechanics. It is a field I have never heard of, and it combines two of my interests - AI and Physics. Through the hackathons, I gained valuable connections and learned a lot from researchers with extensive experience." - Doroteya Stoyanova, Computer Vision Intern
Resources
To ensure you're well-equipped for the Reprogramming AI Models Hackathon, we've compiled a set of resources to support your participation:
- Goodfire's SDK/API with hosted inference: Your primary toolkit for the hackathon. Familiarize yourself with our framework for understanding and modifying AI model behavior.
- Hosted inference on Llama 3 8B and 70B models
- Feature inspection and intervention capabilities: https://docs.goodfire.ai/examples/advanced.html#Feature-intervention-modes
- Example notebooks and tutorials: https://docs.goodfire.ai/examples/quickstart.html#Use-contrastive-features-to-fine-tune-with-a-single-example!
- Latent explorer visualization tools: https://docs.goodfire.ai/examples/latent_explorer.html
- Rate limits and usage guidelines: https://docs.goodfire.ai/rate-limits.html
- Research Preview Playground
- Sandbox environment for model experimentation: https://docs.goodfire.ai/examples/quickstart.html#Replace-model-calls-with-OpenAI-compatible-API
- Feature activation analysis tools: https://docs.goodfire.ai/examples/advanced.html#Feature-intervention-modes
- Conditional intervention testing: https://docs.goodfire.ai/examples/advanced.html#Conditional-feature-interventions
- Check out the Jupyter Notebook Quickstart: . In this quickstart, you'll learn how to:
- Sample from a language model (in this case, Llama 3 8B)
- Search for exciting features and intervene in them to steer the model
- Find features by contrastive search
- Save and load Llama models with steering applied
- Tutorial: Visualizing AI Model Internals: Watch this video to understand how to use Goodfire's tools to map and visualize AI model behavior.
- The Cognitive Revolution Podcast - Episode on Interpretability. n this episode of The Cognitive Revolution, we delve into the science of understanding AI models' inner workings, recent breakthroughs, and the potential impact on AI safety and control
- Auto-interp Paper: This paper applies automation to the problem of scaling an interpretability technique to all the neurons in a large language model.
- ARENA Interpretability with SAEs
- Gemma Scope: a comprehensive, open suite of sparse autoencoders for language model interpretability.
- Neuronpedia: Platform for accelerating research into Sparse Autoencoders
- The Geometry of Concepts: Sparse Autoencoder Feature Structure Paper. This paper investigates the structured organization of concept representations within large language models using sparse autoencoders, revealing a multi-scale structure with refined atomic parallelogram forms, modular brain-like spatial features, and anisotropic galaxy-scale distributions with unique eigenvalue properties.
- Open Source Replication of Anthropic’s Crosscoder paper for model-diffing
- Lesswrong search for SAE
Schedule
Here is the schedule for the Hackathon:
We start with an introductory talk and end the event during the following week with an awards ceremony. Join the public ICal here. You will also find Explorer events, such as collaborative brainstorming and team match-making before the hackathon begins on Discord and in the calendar.
Speakers

Joseph Bloom
Speaker
Joseph co-founded Decode Research, a non-profit organization aiming to accelerate progress in AI safety research infrastructure, and is a mechanistic interpretability researcher.

Callum McDougall
Speaker
ARENA Director and SERI-MATS alumnus specializing in mechanistic interpretability and AI alignment education

Esben Kran
Organizer and Keynote Speaker
Esben is the founder of Apart Research, which he launched at age 22 after leaving grad school. Apart accelerates AI safety research worldwide, producing 20+ papers, award-winning benchmarks like DarkBench, and engaging 4,000+ hackers in research sprints.
Recently co-launched Seldon to fund critical infrastructure for humanity's future, with first investments in Andon Labs, Lucid Computing, Workshop Labs, and Asymmetric Security.
Judges and mentors
Organizers

Tom McGrath
Organizer & Judge

Dan Balsam
Organizer & Mentor

Myra Deng
Organizer
- (opens in new tab)

Archana Vaidheeswaran
Organizer
- (opens in new tab)

Jaime Raldua
Organizer
- (opens in new tab)

Jason Hoelscher-Obermaier
Organizer and Judge
Local sites
AISIG - Decode AI's Black Box & Engineer Model Behavior Hackathon
Join us for the Hackathon in Hereplein 4, 9711GA, Groningen!
Event page: AISIG - Decode AI's Black Box & Engineer Model Behavior Hackathon (opens in new tab)Reprogramming AI Models Hackathon
This is a collaboration between Warwick AI and Warwick Effective Altruism. We will be hosting groups that wish to participate in the hackathon for the weekend.
Event page: Reprogramming AI Models Hackathon (opens in new tab)Reprogramming AI Models Hackathon: CAISH
Cambridge hub for hosting the Reprogramming AI hackathon. Office available with monitors and snacks!
Reprogramming AI Models Hackathon: Edinburgh
Reprogramming AI Models Hackathon: EPFL hub
A local hub for the hackathon on the EPFL campus (luma coming soon). We will provide a room, snacks and drinks for the participants.
Where a Sprint can lead
How our programs connectAnyone can join
Stand out
6 to 16 weeks on your own project, with a research project manager, compute and publication support.
Upcoming Sprints
All SprintsAI Collusion Research Sprint
A weekend research sprint on collusion between AI agents: when it emerges in markets and everyday workflows, how to detect and audit it, how it is carried, and what breaks it. Co-organized with Poseidon Research and AE Studio, online with in-person hubs at Collider in New York City and AI Safety Hong Kong. Top teams are invited to apply to the Apart Fellowship.
Read the brief: AI Collusion Research SprintAI x Epistemics Research Sprint
A weekend research sprint on AI for epistemics: evaluating whether models know how solid their claims are, building trust infrastructure that people and agents can consume, and shipping epistemic products that improve real decisions. Online, four tracks including an open track. Top teams are invited to apply to the Apart Fellowship.
Read the brief: AI x Epistemics Research SprintQuestions? sprints@apartresearch.com




