Classification on Latent Feature Activation for Detecting Adversarial Prompt Vulnerabilities
Hoang-Long Tran, Jack Kaunismaa, Edward Stevinson, Parv Mahajan, Oliver Clive-Griffin · Team Feature Disruption Lab
Submitted to Reprogramming AI Models Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
We present a method leveraging Sparse Autoencoder (SAE)-derived feature activations to identify and mitigate adversarial prompt hijacking in large language models (LLMs). By training a logistic regression classifier on SAE-derived features, we accurately classify diverse adversarial prompts and distinguish between successful and unsuccessful attacks. Utilizing the Goodfire SDK with the LLaMA-8B model, we explored latent feature activations to gain insights into adversarial interactions. This approach highlights the potential of SAE activations for improving LLM safety by enabling automated auditing based on model internals. Future work will focus on scaling this method and exploring its integration as a control mechanism for mitigating attacks.
Reviews
This work covers an important problem and applies a sensible methodology. The performance of the results is impressive - I had to check in the code that the results were in fact on a test set. I'd be interested in seeing how often harmless prompts are misclassified though. Definitely worth extending further - these results are quite promising.
This is addressing a really important practical problem. I liked their approach, am impressed with the results, and would be keen to see them build on this work!
Cite this project
@misc{tran2024classification,
title = {{Classification on Latent Feature Activation for Detecting Adversarial Prompt Vulnerabilities}},
author = {Hoang-Long Tran and Jack Kaunismaa and Edward Stevinson and Parv Mahajan and Oliver Clive-Griffin},
year = {2024},
month = nov,
note = {Submitted to Reprogramming AI Models Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/classification-on-latent-feature-activation-for-detecting-adversarial-prompt-vulnerabilities}},
url = {https://apartresearch.com/sprints/projects/classification-on-latent-feature-activation-for-detecting-adversarial-prompt-vulnerabilities}
}More from Reprogramming AI Models Hackathon
- 1st place by peer reviewView project: AutoSteer: Weight-Preserving Reinforcement Learning for Interpretable Model Control
AutoSteer: Weight-Preserving Reinforcement Learning for Interpretable Model Control
AI Safety Initiative Groningen
Traditional fine-tuning methods for language models, while effective, often disrupt internal model features that could provide valuable insights into model behavior. We present a novel approach combining Reinforcement …
- View project: Utilitarian Decision-Making in Models - Evaluation and Steering
Utilitarian Decision-Making in Models - Evaluation and Steering
Byte To Brain
We design an eval based on the Oxford Utilitarianism Scale (OUS) that measures the model’s deontological vs utilitarian preference in a nine question, two factor model. Using this scale, we measure how feature steering …
- View project: Steering Swiftly to Safety with Sparse Autoencoders
Steering Swiftly to Safety with Sparse Autoencoders
Explaining_Polysemantic_Feature_Learning
We explore using SAEs for unlearning dangerous capabilities in a cheaper and more interpretable way.