Detecting Deception with AI Tics 😉
Samuel Svenningsen, Ilan Moscovitz, Nikhil Kotecha · Team Polygraph
Submitted to Deception Detection Hackathon: Preventing AI deception. Sprint projects are early-stage work by participants, not Apart Research publications.
We present a novel approach: intentionally inducing subtle "tics" in AI responses as a marker for deceptive behavior. By adding a system prompt, we embed innocuous yet detectable patterns that manifest when the AI knowingly engages in deception.
Reviews
This work tests whether models follow a system prompt that tells them to mark any lie with a winky face. The robustness of the method is not discussed (though fine-tuning it in is suggested), nor is it compared to existing deception detection methods (e.g. contrast-consistent search and variants, activation vectors, blackbox lie detection). It is suggested that future models would find it hard to not include such tells if fine-tuned, but this statement is not supported. In particular, deception is only a problem if fine-tuning models for helpfulness & harmlessness does not work, so in any deception scenario, simple at least some simple fine-tuning strategies have already broken. Future work could explore whether fine-tuning such tells are more robust than other types of fine-tuning aimed at preventing or detecting deception.
Cite this project
@misc{svenningsen2024detecting,
title = {{Detecting Deception with AI Tics 😉}},
author = {Samuel Svenningsen and Ilan Moscovitz and Nikhil Kotecha},
year = {2024},
month = jul,
note = {Submitted to Deception Detection Hackathon: Preventing AI deception, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/detecting-deception-with-ai-tics}},
url = {https://apartresearch.com/sprints/projects/detecting-deception-with-ai-tics}
}More from Deception Detection Hackathon: Preventing AI deception
- 1st place by peer reviewView project: Sandbag Detection through Model Degradation
Sandbag Detection through Model Degradation
Team Truth Serum
We propose a novel technique to detect sandbagging in LLMs by adding varying amount of noise to model weights and monitoring performance.
- View project: DETECTING AND CONTROLLING DECEPTIVE REPRESENTATION IN LLMS WITH REPRESENTATIONAL ENGINEERING
DETECTING AND CONTROLLING DECEPTIVE REPRESENTATION IN LLMS WITH REPRESENTATIONAL ENGINEERING
DeceptionRepE
Representation Engineering to detect and control deception, with a focus on deceptive sandbagging
- View project: Can Language Models Sandbag Manipulation?
Can Language Models Sandbag Manipulation?
Can Language Models Sandbag Manipulation?
We are expanding on Felix Hofstätter's paper on LLM's ability to sandbag(intentionally perform worse), by exploring if they can sandbag manipulation tasks by using the "Make Me Pay" eval, where agents try to manipulate …