Skip to content
Sprint projectJul 1, 2024

Evaluating Steering Methods for Deceptive Behavior Control in LLMs

Casey Hird, Basavasagar Patil, Tinuade Adeleke, Adam Fraknoi, Neel Jay · Team Deception Detecion Ninjas

Submitted to Deception Detection Hackathon: Preventing AI deception. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Evaluating Steering Methods for Deceptive Behavior Control in LLMs

Share

We use SOTA steering methods, including CAA, LAT, and SAEs to find and control deceptive behaviors in LLMs. We also release a new deception dataset, and demonstrate that the dataset and the prompt formatting used are significant when evaluating the efficacy of steering methods.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

  1. It’s great to see the paper give such a good  overview of how the different methods (CAA/SAE/LAT) can be used for evaluating deception. The limitations identified are valid and I’d be interested to see how the exiting work builds to address them, particularly having more explicit comparison of the effects of each method and expanding the analysis to more models would be very valuable.

  2. An in-depth review of control vectors for deception mitigation over SAEs, LATs, and CAA. Great overview of the various methods’ effects depending on hyperparameter tuning. One potential extension of the work could be qualitative tendency analysis of the resulting outputs using each vector. E.g. when SAEs were used for Golden Gate Claude, individuals with OCD mentioned it was similar to their experience. Might CAAs and LATs just “feel” different than SAEs? Or would we expect them to have a similar effect? The statistical results are of course very solid. Great approach to those and great work overall!

Cite this project

@misc{hird2024evaluating,
  title = {{Evaluating Steering Methods for Deceptive Behavior Control in LLMs}},
  author = {Casey Hird and Basavasagar Patil and Tinuade Adeleke and Adam Fraknoi and Neel Jay},
  year = {2024},
  month = jul,
  note = {Submitted to Deception Detection Hackathon: Preventing AI deception, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/evaluating-steering-methods-for-deceptive-behavior-control-in-llms}},
  url = {https://apartresearch.com/sprints/projects/evaluating-steering-methods-for-deceptive-behavior-control-in-llms}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026