Detection of potentially deceptive attitudes using expression style analysis
Roland Pihlakas · Team ClarityNaut
Submitted to Deception Detection Hackathon: Preventing AI deception. Sprint projects are early-stage work by participants, not Apart Research publications.
My work on this hackathon consists of two parts: 1) As a sanity check, verifying the deception execution capability of GPT4. The conclusion is “definitely yes”. I provide a few arguments about when that is a useful functionality. 2) Experimenting with recognising potential deception by using an LLM-based text analysis algorithm to highlight certain manipulative expression styles sometimes present in the deceptive responses. For that task I pre-selected a small subset of input data consisting only of entries containing responses with elements of psychological influence. The results show that LLM-based text analysis is able to detect different manipulative styles in responses, or alternatively, attitudes leading to deception in case of internal thoughts.
Reviews
No public critique yet.
Cite this project
@misc{pihlakas2024detection,
title = {{Detection of potentially deceptive attitudes using expression style analysis}},
author = {Roland Pihlakas},
year = {2024},
month = jul,
note = {Submitted to Deception Detection Hackathon: Preventing AI deception, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/detection-of-potentially-deceptive-attitudes-using-expression-style-analysis}},
url = {https://apartresearch.com/sprints/projects/detection-of-potentially-deceptive-attitudes-using-expression-style-analysis}
}More from Deception Detection Hackathon: Preventing AI deception
- 1st place by peer reviewView project: Sandbag Detection through Model Degradation
Sandbag Detection through Model Degradation
Team Truth Serum
We propose a novel technique to detect sandbagging in LLMs by adding varying amount of noise to model weights and monitoring performance.
- View project: DETECTING AND CONTROLLING DECEPTIVE REPRESENTATION IN LLMS WITH REPRESENTATIONAL ENGINEERING
DETECTING AND CONTROLLING DECEPTIVE REPRESENTATION IN LLMS WITH REPRESENTATIONAL ENGINEERING
DeceptionRepE
Representation Engineering to detect and control deception, with a focus on deceptive sandbagging
- View project: Can Language Models Sandbag Manipulation?
Can Language Models Sandbag Manipulation?
Can Language Models Sandbag Manipulation?
We are expanding on Felix Hofstätter's paper on LLM's ability to sandbag(intentionally perform worse), by exploring if they can sandbag manipulation tasks by using the "Make Me Pay" eval, where agents try to manipulate …