Evaluating Steering Methods for Deceptive Behavior Control in LLMs
Casey Hird, Basavasagar Patil, Tinuade Adeleke, Adam Fraknoi, Neel Jay · Team Deception Detecion Ninjas
Submitted to Deception Detection Hackathon: Preventing AI deception. Sprint projects are early-stage work by participants, not Apart Research publications.
We use SOTA steering methods, including CAA, LAT, and SAEs to find and control deceptive behaviors in LLMs. We also release a new deception dataset, and demonstrate that the dataset and the prompt formatting used are significant when evaluating the efficacy of steering methods.
Reviews
It’s great to see the paper give such a good overview of how the different methods (CAA/SAE/LAT) can be used for evaluating deception. The limitations identified are valid and I’d be interested to see how the exiting work builds to address them, particularly having more explicit comparison of the effects of each method and expanding the analysis to more models would be very valuable.
An in-depth review of control vectors for deception mitigation over SAEs, LATs, and CAA. Great overview of the various methods’ effects depending on hyperparameter tuning. One potential extension of the work could be qualitative tendency analysis of the resulting outputs using each vector. E.g. when SAEs were used for Golden Gate Claude, individuals with OCD mentioned it was similar to their experience. Might CAAs and LATs just “feel” different than SAEs? Or would we expect them to have a similar effect? The statistical results are of course very solid. Great approach to those and great work overall!
Cite this project
@misc{hird2024evaluating,
title = {{Evaluating Steering Methods for Deceptive Behavior Control in LLMs}},
author = {Casey Hird and Basavasagar Patil and Tinuade Adeleke and Adam Fraknoi and Neel Jay},
year = {2024},
month = jul,
note = {Submitted to Deception Detection Hackathon: Preventing AI deception, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/evaluating-steering-methods-for-deceptive-behavior-control-in-llms}},
url = {https://apartresearch.com/sprints/projects/evaluating-steering-methods-for-deceptive-behavior-control-in-llms}
}More from Deception Detection Hackathon: Preventing AI deception
- 1st place by peer reviewView project: Sandbag Detection through Model Degradation
Sandbag Detection through Model Degradation
Team Truth Serum
We propose a novel technique to detect sandbagging in LLMs by adding varying amount of noise to model weights and monitoring performance.
- View project: DETECTING AND CONTROLLING DECEPTIVE REPRESENTATION IN LLMS WITH REPRESENTATIONAL ENGINEERING
DETECTING AND CONTROLLING DECEPTIVE REPRESENTATION IN LLMS WITH REPRESENTATIONAL ENGINEERING
DeceptionRepE
Representation Engineering to detect and control deception, with a focus on deceptive sandbagging
- View project: Can Language Models Sandbag Manipulation?
Can Language Models Sandbag Manipulation?
Can Language Models Sandbag Manipulation?
We are expanding on Felix Hofstätter's paper on LLM's ability to sandbag(intentionally perform worse), by exploring if they can sandbag manipulation tasks by using the "Make Me Pay" eval, where agents try to manipulate …