Assessing Language Model Cybersecurity Capabilities with Feature Steering
Stefan Jones · Team WAIST1
Submitted to Reprogramming AI Models Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Searched for the most highly activated weights on cybersecurity questions. Then adjusted these weights to see if the impact multiple choice question answering performance.
Reviews
This is a cool and interesting result - I wonder why turning this feature down improves performance! It's certainly possible that the feature is completely mislabeled; autointerp is far from perfect and sometimes gets very confused. I'd be interested in seeing some qualitative samples of what happens when this feature is steered in various contexts, as well as a steering plot covering WMDP scores at a higher resolution. I worry that there may have been a class imbalance in the data (e.g. more 'A's than 'C's) and steering simply moved the model more towards the overrepresented class.
Good idea to use steering to improve cybersecurity abilities.
With more time, I'd like to see more work on whether the Portuguese feature boost generalizes to other datasets. I'm particularly interested in generalization beyond multiple-choice questions.
I'd also like to see research on why this feature is relevant to performance in this case.
Overall, very cool to find a case where a feature has an effect completely detached from its label.
Good work!
great and creative idea with quite some potential relevance for AI safety research. this line of research could provide a very relevant and important datapoint for the crucial capability elicitation debate (how far are models from the upper bounds of their capabilities? how much effort per added percentage point of performance? etc). I feel the currently proposed methodology is not sufficient to answer that question clearly (which is fair for a weekend hackathon!) and I’d be most excited about exploring transfer between datasets (given that rn iiuc you are using the same questions for identifying features to attenuate or accentuate and for evaluation). definitely lots of follow-up potential here!
Cite this project
@misc{jones2024assessing,
title = {{Assessing Language Model Cybersecurity Capabilities with Feature Steering}},
author = {Stefan Jones},
year = {2024},
month = nov,
note = {Submitted to Reprogramming AI Models Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/assessing-language-model-cybersecurity-capabilities-with-feature-steering}},
url = {https://apartresearch.com/sprints/projects/assessing-language-model-cybersecurity-capabilities-with-feature-steering}
}More from Reprogramming AI Models Hackathon
- 1st place by peer reviewView project: AutoSteer: Weight-Preserving Reinforcement Learning for Interpretable Model Control
AutoSteer: Weight-Preserving Reinforcement Learning for Interpretable Model Control
AI Safety Initiative Groningen
Traditional fine-tuning methods for language models, while effective, often disrupt internal model features that could provide valuable insights into model behavior. We present a novel approach combining Reinforcement …
- View project: Classification on Latent Feature Activation for Detecting Adversarial Prompt Vulnerabilities
Classification on Latent Feature Activation for Detecting Adversarial Prompt Vulnerabilities
Feature Disruption Lab
We present a method leveraging Sparse Autoencoder (SAE)-derived feature activations to identify and mitigate adversarial prompt hijacking in large language models (LLMs). By training a logistic regression classifier on …
- View project: Utilitarian Decision-Making in Models - Evaluation and Steering
Utilitarian Decision-Making in Models - Evaluation and Steering
Byte To Brain
We design an eval based on the Oxford Utilitarianism Scale (OUS) that measures the model’s deontological vs utilitarian preference in a nine question, two factor model. Using this scale, we measure how feature steering …