Feature based unlearning
Patrick Quinn, Yucheng Sun · Team Model memory analysis
Submitted to Reprogramming AI Models Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
An exploration of using features to perform unlearning on answering trivia questions.
Reviews
A nice, practical approach to unlearning via sparse autoencoders! The features often being related to general question answering demonstrates one of the important challenges with scaling unlearning generally. I agree with the point about analysis of those manually discovered, more effective features being useful and it'd be cool to see if there's some sort of automated LLM workflow that would be able to surface those same features with less effort.
This is a good attempt at answering an important safety-relevant question. Unfortunately the current setup doesn't work well enough to accurately ablate factual knowledge, but it was worth trying and the methodology used here is sufficient to answer the question.
It's possible that using attribution would have improved feature selection and made it more automated. The results in table 1 are impressive - I'd be interested in seeing more failure cases however as (as the authors indicate) Figure 1 tells a different story at the scale of the entire dataset.
The results are pretty impresive given the time constraints and the API rate limits. I would love to see an extension of this work with smaller models and more features
Cite this project
@misc{quinn2024feature,
title = {{Feature based unlearning}},
author = {Patrick Quinn and Yucheng Sun},
year = {2024},
month = nov,
note = {Submitted to Reprogramming AI Models Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/feature-based-unlearning}},
url = {https://apartresearch.com/sprints/projects/feature-based-unlearning}
}More from Reprogramming AI Models Hackathon
- 1st place by peer reviewView project: AutoSteer: Weight-Preserving Reinforcement Learning for Interpretable Model Control
AutoSteer: Weight-Preserving Reinforcement Learning for Interpretable Model Control
AI Safety Initiative Groningen
Traditional fine-tuning methods for language models, while effective, often disrupt internal model features that could provide valuable insights into model behavior. We present a novel approach combining Reinforcement …
- View project: Classification on Latent Feature Activation for Detecting Adversarial Prompt Vulnerabilities
Classification on Latent Feature Activation for Detecting Adversarial Prompt Vulnerabilities
Feature Disruption Lab
We present a method leveraging Sparse Autoencoder (SAE)-derived feature activations to identify and mitigate adversarial prompt hijacking in large language models (LLMs). By training a logistic regression classifier on …
- View project: Utilitarian Decision-Making in Models - Evaluation and Steering
Utilitarian Decision-Making in Models - Evaluation and Steering
Byte To Brain
We design an eval based on the Oxford Utilitarianism Scale (OUS) that measures the model’s deontological vs utilitarian preference in a nine question, two factor model. Using this scale, we measure how feature steering …