Skip to content
Sprint projectNov 25, 2024

Assessing Language Model Cybersecurity Capabilities with Feature Steering

Stefan Jones · Team WAIST1

Submitted to Reprogramming AI Models Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Assessing Language Model Cybersecurity Capabilities with Feature Steering

Code (opens in new tab)
Share

Searched for the most highly activated weights on cybersecurity questions. Then adjusted these weights to see if the impact multiple choice question answering performance.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

Does the project contribute to the field of mechanistic interpretability? Does it provide new insights into understanding or steering AI model behavior? How well does it move us towards reprogramming AI models? How original and innovative is the approach?

How important is the contribution to advancing the field of AI safety? Do we expect the results to generalize beyond the specific case(s) presented in the submission? Does the approach introduce new safety mechanisms or enhance existing ones in innovative ways?

How well is the project executed from a technical standpoint?, Is the code well-structured, documented, and reproducible?, How effectively does it utilize Goodfire's SDK/API and other provided resources?, How clearly and effectively is the research presented in the paper?, Quality of visualizations and demos (if applicable), Clarity of methodology explanation and results interpretation

  1. This is a cool and interesting result - I wonder why turning this feature down improves performance! It's certainly possible that the feature is completely mislabeled; autointerp is far from perfect and sometimes gets very confused. I'd be interested in seeing some qualitative samples of what happens when this feature is steered in various contexts, as well as a steering plot covering WMDP scores at a higher resolution. I worry that there may have been a class imbalance in the data (e.g. more 'A's than 'C's) and steering simply moved the model more towards the overrepresented class.

  2. Good idea to use steering to improve cybersecurity abilities.

    With more time, I'd like to see more work on whether the Portuguese feature boost generalizes to other datasets. I'm particularly interested in generalization beyond multiple-choice questions.

    I'd also like to see research on why this feature is relevant to performance in this case.

    Overall, very cool to find a case where a feature has an effect completely detached from its label.

    Good work!

  3. great and creative idea with quite some potential relevance for AI safety research. this line of research could provide a very relevant and important datapoint for the crucial capability elicitation debate (how far are models from the upper bounds of their capabilities? how much effort per added percentage point of performance? etc). I feel the currently proposed methodology is not sufficient to answer that question clearly (which is fair for a weekend hackathon!) and I’d be most excited about exploring transfer between datasets (given that rn iiuc you are using the same questions for identifying features to attenuate or accentuate and for evaluation). definitely lots of follow-up potential here!

Cite this project

@misc{jones2024assessing,
  title = {{Assessing Language Model Cybersecurity Capabilities with Feature Steering}},
  author = {Stefan Jones},
  year = {2024},
  month = nov,
  note = {Submitted to Reprogramming AI Models Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/assessing-language-model-cybersecurity-capabilities-with-feature-steering}},
  url = {https://apartresearch.com/sprints/projects/assessing-language-model-cybersecurity-capabilities-with-feature-steering}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026