Skip to content
Sprint projectNov 24, 2024

Steering Swiftly to Safety with Sparse Autoencoders

Agatha Duzan, Guillaume Martres, Syrine Noame, Abhinand Shibu, Flavia Wallenhorst, Arthur Wuhrmann · Team Explaining_Polysemantic_Feature_Learning

Submitted to Reprogramming AI Models Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Steering Swiftly to Safety with Sparse Autoencoders

Code (opens in new tab)
Share

We explore using SAEs for unlearning dangerous capabilities in a cheaper and more interpretable way.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

Does the project contribute to the field of mechanistic interpretability? Does it provide new insights into understanding or steering AI model behavior? How well does it move us towards reprogramming AI models? How original and innovative is the approach?

How important is the contribution to advancing the field of AI safety? Do we expect the results to generalize beyond the specific case(s) presented in the submission? Does the approach introduce new safety mechanisms or enhance existing ones in innovative ways?

How well is the project executed from a technical standpoint?, Is the code well-structured, documented, and reproducible?, How effectively does it utilize Goodfire's SDK/API and other provided resources?, How clearly and effectively is the research presented in the paper?, Quality of visualizations and demos (if applicable), Clarity of methodology explanation and results interpretation

  1. I like the comparison between different approaches to using SAEs to unlearn dangerous knowledge. This project uses a very sensible approach (e.g. including MMLU as standard performance benchmark) and is clearly presented. In future, it could be interesting to explore the robustness of the unlearning and why performance on MMLU comp sci appears to increase.

  2. Very well structured work and a very insightful comparison. Pity that the Goodfire API restrictions got in the way of reproducing the reference study. The proposed extension about testing robustness seems very promising and worth exploring.

Cite this project

@misc{duzan2024steering,
  title = {{Steering Swiftly to Safety with Sparse Autoencoders}},
  author = {Agatha Duzan and Guillaume Martres and Syrine Noame and Abhinand Shibu and Flavia Wallenhorst and Arthur Wuhrmann},
  year = {2024},
  month = nov,
  note = {Submitted to Reprogramming AI Models Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/steering-swiftly-to-safety-with-sparse-autoencoders}},
  url = {https://apartresearch.com/sprints/projects/steering-swiftly-to-safety-with-sparse-autoencoders}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026