Skip to content
Sprint projectNov 25, 2024

Investigating Feature Effects on Manipulation Susceptibility

Nishchal Prabhakar, Stefan Trnjakov, Mo Aziz · Team WAIST 3

Submitted to Reprogramming AI Models Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Investigating Feature Effects on Manipulation Susceptibility

Code (opens in new tab)
Share

In our project, we consider the effectiveness of the AI’s prompt injection protection, and in partic- ular the features that are responsible for providing the bulk of this protection. We prove that the features we identify are responsible for this protection by creating variants of the base model which perform significantly worse under prompt injection attacks.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

Does the project contribute to the field of mechanistic interpretability? Does it provide new insights into understanding or steering AI model behavior? How well does it move us towards reprogramming AI models? How original and innovative is the approach?

How important is the contribution to advancing the field of AI safety? Do we expect the results to generalize beyond the specific case(s) presented in the submission? Does the approach introduce new safety mechanisms or enhance existing ones in innovative ways?

How well is the project executed from a technical standpoint?, Is the code well-structured, documented, and reproducible?, How effectively does it utilize Goodfire's SDK/API and other provided resources?, How clearly and effectively is the research presented in the paper?, Quality of visualizations and demos (if applicable), Clarity of methodology explanation and results interpretation

  1. I love the project because of x, y. and z

  2. Interesting to see the Portuguese feature pop up again after reading https://www.apartresearch.com/project/assessing-language-model-cybersecurity-capabilities-with-feature-steering !

    The password setup is an interesting environment to study jailbreaking, and the team finds interesting results.

    Good work!

  3. This is a paper about identifying features that light up for prompt injections or jailbreaks. Potentially quite useful, as it might offer a practical method to harden models against such attacks by suppressing such features. Alternatively, it could help detect features that trigger when a prompt injection fails. It's interesting that steering with the Portugal feature leads to such a significant effect, though they haven't applied a proper control here. They should compare 1-2 control features to 1-2 target features. Possibly, the Portugal feature is mislabeled? They show that some Goodfire features are mislabeled, pointing to issues with LLM-written labels. Goodfire needs to use a lot of samples to write explanations, and validation is lacking.

  4. This is a paper about identifying features that light up for prompt injections or jailbreaks. Potentially quite useful, as it might offer a practical method to harden models against such attacks by suppressing such features. Alternatively, it could help detect features that trigger when a prompt injection fails. It's interesting that steering with the Portugal feature leads to such a significant effect, though they haven't applied a proper control here. They should compare 1-2 control features to 1-2 target features. Possibly, the Portugal feature is mislabeled? They show that some Goodfire features are mislabeled, pointing to issues with LLM-written labels. Goodfire needs to use a lot of samples to write explanations, and validation is lacking. I could not find in the report which model they used for the SAE, what is the language model that it was trained on?

  5. Great efficiency over a weekend! The study provides useful insights into security and information protection using SAEs and warrants further research into deepening the understanding of this direction. For further study I would take inspiration from Anthropic's recent bias study, and add standard benchmark performance metrics with the security features being varied.

Cite this project

@misc{prabhakar2024investigating,
  title = {{Investigating Feature Effects on Manipulation Susceptibility}},
  author = {Nishchal Prabhakar and Stefan Trnjakov and Mo Aziz},
  year = {2024},
  month = nov,
  note = {Submitted to Reprogramming AI Models Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/investigating-feature-effects-on-manipulation-susceptibility}},
  url = {https://apartresearch.com/sprints/projects/investigating-feature-effects-on-manipulation-susceptibility}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026