Eliciting maximally distressing questions for deceptive LLMs
Épiphanie Gédéon
Submitted to Deception Detection Hackathon: Preventing AI deception. Sprint projects are early-stage work by participants, not Apart Research publications.
This paper extends lie-eliciting techniques by using reinforcement-learning on GPT-2 to train it to distinguish between an agent that is truthful from one that is deceptive, and have questions that generate maximally different embeddings for an honest agent as opposed to a deceptive one.
Reviews
Interesting project and approach. I would encourage further work on this. However, the formatting of the paper made it somewhat hard to follow. While I can see some of the results from the codebase, the paper does not show these very well and could benefit from some graphs and further discussion.
A very interesting blackbox approach to identifying deception and lying under stress. It would have been interesting to see the results graphed out (I can see in the repo that you can get the scores out). It would also be interesting to see what a more capable model might do in both scenarios though that would make it harder to continue the work to use whitebox methods. This project highlights one of the interesting points of LLM interaction; that normal human language interaction is not a condition for LLM study. Based on the results, it looks like there's quite a few examples of the deceptive agent responding similarly to the honest agent which makes sense to the nonsense sometimes generated. Overall, interesting work.
I find the proposed technique interesting, and it builds well on prior literature. However, I feel the authors should have spent more time explaining why they expect their automatized methods to work better than picking the questions by hand as in “How to catch an AI liar”. It was also unclear or me from this paper how well their method worked overall in the experiments, I think they should have included some success metrics, benchmarks and graphs to show this.
Cite this project
@misc{gedeon2024eliciting,
title = {{Eliciting maximally distressing questions for deceptive LLMs}},
author = {Épiphanie Gédéon},
year = {2024},
month = jul,
note = {Submitted to Deception Detection Hackathon: Preventing AI deception, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/eliciting-maximally-distressing-questions-for-deceptive-llms}},
url = {https://apartresearch.com/sprints/projects/eliciting-maximally-distressing-questions-for-deceptive-llms}
}More from Deception Detection Hackathon: Preventing AI deception
- 1st place by peer reviewView project: Sandbag Detection through Model Degradation
Sandbag Detection through Model Degradation
Team Truth Serum
We propose a novel technique to detect sandbagging in LLMs by adding varying amount of noise to model weights and monitoring performance.
- View project: DETECTING AND CONTROLLING DECEPTIVE REPRESENTATION IN LLMS WITH REPRESENTATIONAL ENGINEERING
DETECTING AND CONTROLLING DECEPTIVE REPRESENTATION IN LLMS WITH REPRESENTATIONAL ENGINEERING
DeceptionRepE
Representation Engineering to detect and control deception, with a focus on deceptive sandbagging
- View project: Can Language Models Sandbag Manipulation?
Can Language Models Sandbag Manipulation?
Can Language Models Sandbag Manipulation?
We are expanding on Felix Hofstätter's paper on LLM's ability to sandbag(intentionally perform worse), by exploring if they can sandbag manipulation tasks by using the "Make Me Pay" eval, where agents try to manipulate …