DETECTING AND CONTROLLING DECEPTIVE REPRESENTATION IN LLMS WITH REPRESENTATIONAL ENGINEERING
Avyay M Casheekar, Kaushik Sanjay Prabhakar, Kanishk Rath, Sienka Dounia · Team DeceptionRepE
Submitted to Deception Detection Hackathon: Preventing AI deception. Sprint projects are early-stage work by participants, not Apart Research publications.
Representation Engineering to detect and control deception, with a focus on deceptive sandbagging
Reviews
Very well written, explained and excellent visualisations. It might also be informative to have more contextual analysis of where the various methods fail. Would be curious to see you look into why the Positive Deceptive Control might be increasing the amount of deceptive output in samples where the initial prompt was designed to elicit truthfulness.
The experiments are well executed, though the results could be presented somewhat more clearly. I think it would be better however to make it more clear what probes and similar techniques can a cannot do, and how do they compare to baseline techniques like finetuning and RLHF.
Cite this project
@misc{casheekar2024detecting,
title = {{DETECTING AND CONTROLLING DECEPTIVE REPRESENTATION IN LLMS WITH REPRESENTATIONAL ENGINEERING}},
author = {Avyay M Casheekar and Kaushik Sanjay Prabhakar and Kanishk Rath and Sienka Dounia},
year = {2024},
month = jul,
note = {Submitted to Deception Detection Hackathon: Preventing AI deception, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/detecting-and-controlling-deceptive-representation-in-llms-with-representational-engineering}},
url = {https://apartresearch.com/sprints/projects/detecting-and-controlling-deceptive-representation-in-llms-with-representational-engineering}
}More from Deception Detection Hackathon: Preventing AI deception
- 1st place by peer reviewView project: Sandbag Detection through Model Degradation
Sandbag Detection through Model Degradation
Team Truth Serum
We propose a novel technique to detect sandbagging in LLMs by adding varying amount of noise to model weights and monitoring performance.
- View project: Can Language Models Sandbag Manipulation?
Can Language Models Sandbag Manipulation?
Can Language Models Sandbag Manipulation?
We are expanding on Felix Hofstätter's paper on LLM's ability to sandbag(intentionally perform worse), by exploring if they can sandbag manipulation tasks by using the "Make Me Pay" eval, where agents try to manipulate …
- View project: Deceptive behavior does not seem to be reducible to a single vector
Deceptive behavior does not seem to be reducible to a single vector
Team Whitebox Research
Given the success of previous work in showing that complex behaviors in LLMs can be mediated by single directions or vectors, such as refusing to answer potentially harmful prompts, we investigated if we can create a …