Sep 15, 2024Online and in person
ARENA 4.0 Interpretability Hackathon
Hackathon for working on interpretability projects during ARENA v4. These could be training or interpreting SAEs, finding circuits in GPT2-Small, training & interpreting toy transformer models on algorithmic tasks, building visualization software / infra for interpretability research, or anything else you can think of!
Entries
- View project: Interpreting a toy model for finding the maximum element in a list
Interpreting a toy model for finding the maximum element in a list
Interpreting a toy model for finding the maximum element in a list
- View project: nnsight transparent debugging
nnsight transparent debugging
We started this project with the intent of identifying a specific issue with nnsight debugging and submitting a pull request to fix it. We found a minimal test case where an IndexError within a nnsight run wasn’t correctly propagated to the user, making debugging difficult, and wrote up a proposal for some pull …
- View project: minTranscoders
minTranscoders
Attempting to be a minGPT like implementation for transcoders for MLP hidden state in transformers - part of ARENA 4.0 Interpretability Hackathon via Apart Research
- View project: Latent Space Clustering and Summarization
Latent Space Clustering and Summarization
I wanted to see how modern dimensionality reduction and clustering approaches can support visualization and interpretation of LLM latent spaces. I explored a number of different approaches and algoriths, but ultimately converged on UMAP for dimensionality reduction and birch clustering to extract groups of tokens in …
Overview
Organized by ARENA
Hackathon for working on interpretability projects during ARENA v4. These could be training or interpreting SAEs, finding circuits in GPT2-Small, training & interpreting toy transformer models on algorithmic tasks, building visualization software / infra for interpretability research, or anything else you can think of!
This event ran on the 15th of September 2024
Where a Sprint can lead
How our programs connectAnyone can join
Stand out
6 to 16 weeks on your own project, with a research project manager, compute and publication support.
Upcoming Sprints
All SprintsAI Collusion Research Sprint
A weekend research sprint on collusion between AI agents: when it emerges in markets and everyday workflows, how to detect and audit it, how it is carried, and what breaks it. Co-organized with Poseidon Research and AE Studio, online with in-person hubs at Collider in New York City and AI Safety Hong Kong. Top teams are invited to apply to the Apart Fellowship.
Read the brief: AI Collusion Research SprintAI x Epistemics Research Sprint
A weekend research sprint on AI for epistemics: evaluating whether models know how solid their claims are, building trust infrastructure that people and agents can consume, and shipping epistemic products that improve real decisions. Online, four tracks including an open track. Top teams are invited to apply to the Apart Fellowship.
Read the brief: AI x Epistemics Research SprintQuestions? sprints@apartresearch.com