Steering Swiftly to Safety with Sparse Autoencoders
Agatha Duzan, Guillaume Martres, Syrine Noame, Abhinand Shibu, Flavia Wallenhorst, Arthur Wuhrmann · Team Explaining_Polysemantic_Feature_Learning
Submitted to Reprogramming AI Models Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
We explore using SAEs for unlearning dangerous capabilities in a cheaper and more interpretable way.
Reviews
I like the comparison between different approaches to using SAEs to unlearn dangerous knowledge. This project uses a very sensible approach (e.g. including MMLU as standard performance benchmark) and is clearly presented. In future, it could be interesting to explore the robustness of the unlearning and why performance on MMLU comp sci appears to increase.
Very well structured work and a very insightful comparison. Pity that the Goodfire API restrictions got in the way of reproducing the reference study. The proposed extension about testing robustness seems very promising and worth exploring.
Cite this project
@misc{duzan2024steering,
title = {{Steering Swiftly to Safety with Sparse Autoencoders}},
author = {Agatha Duzan and Guillaume Martres and Syrine Noame and Abhinand Shibu and Flavia Wallenhorst and Arthur Wuhrmann},
year = {2024},
month = nov,
note = {Submitted to Reprogramming AI Models Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/steering-swiftly-to-safety-with-sparse-autoencoders}},
url = {https://apartresearch.com/sprints/projects/steering-swiftly-to-safety-with-sparse-autoencoders}
}More from Reprogramming AI Models Hackathon
- 1st place by peer reviewView project: AutoSteer: Weight-Preserving Reinforcement Learning for Interpretable Model Control
AutoSteer: Weight-Preserving Reinforcement Learning for Interpretable Model Control
AI Safety Initiative Groningen
Traditional fine-tuning methods for language models, while effective, often disrupt internal model features that could provide valuable insights into model behavior. We present a novel approach combining Reinforcement …
- View project: Classification on Latent Feature Activation for Detecting Adversarial Prompt Vulnerabilities
Classification on Latent Feature Activation for Detecting Adversarial Prompt Vulnerabilities
Feature Disruption Lab
We present a method leveraging Sparse Autoencoder (SAE)-derived feature activations to identify and mitigate adversarial prompt hijacking in large language models (LLMs). By training a logistic regression classifier on …
- View project: Utilitarian Decision-Making in Models - Evaluation and Steering
Utilitarian Decision-Making in Models - Evaluation and Steering
Byte To Brain
We design an eval based on the Oxford Utilitarianism Scale (OUS) that measures the model’s deontological vs utilitarian preference in a nine question, two factor model. Using this scale, we measure how feature steering …