Skip to content
Sprint projectNov 24, 2024

Feature based unlearning

Patrick Quinn, Yucheng Sun · Team Model memory analysis

Submitted to Reprogramming AI Models Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

An exploration of using features to perform unlearning on answering trivia questions.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

Does the project contribute to the field of mechanistic interpretability? Does it provide new insights into understanding or steering AI model behavior? How well does it move us towards reprogramming AI models? How original and innovative is the approach?

How important is the contribution to advancing the field of AI safety? Do we expect the results to generalize beyond the specific case(s) presented in the submission? Does the approach introduce new safety mechanisms or enhance existing ones in innovative ways?

How well is the project executed from a technical standpoint?, Is the code well-structured, documented, and reproducible?, How effectively does it utilize Goodfire's SDK/API and other provided resources?, How clearly and effectively is the research presented in the paper?, Quality of visualizations and demos (if applicable), Clarity of methodology explanation and results interpretation

  1. A nice, practical approach to unlearning via sparse autoencoders! The features often being related to general question answering demonstrates one of the important challenges with scaling unlearning generally. I agree with the point about analysis of those manually discovered, more effective features being useful and it'd be cool to see if there's some sort of automated LLM workflow that would be able to surface those same features with less effort.

  2. This is a good attempt at answering an important safety-relevant question. Unfortunately the current setup doesn't work well enough to accurately ablate factual knowledge, but it was worth trying and the methodology used here is sufficient to answer the question.

    It's possible that using attribution would have improved feature selection and made it more automated. The results in table 1 are impressive - I'd be interested in seeing more failure cases however as (as the authors indicate) Figure 1 tells a different story at the scale of the entire dataset.

  3. The results are pretty impresive given the time constraints and the API rate limits. I would love to see an extension of this work with smaller models and more features

Cite this project

@misc{quinn2024feature,
  title = {{Feature based unlearning}},
  author = {Patrick Quinn and Yucheng Sun},
  year = {2024},
  month = nov,
  note = {Submitted to Reprogramming AI Models Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/feature-based-unlearning}},
  url = {https://apartresearch.com/sprints/projects/feature-based-unlearning}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026