Boosting Language Model Honesty with Truthful Suffixes
Smitty van Bodegom, Giles Edkins, Annie Szorkin · Team Honest Algorithms, Eh
Submitted to Deception Detection Hackathon: Preventing AI deception. Sprint projects are early-stage work by participants, not Apart Research publications.
We investigate the construction of truthful suffixes, which cause models to provide more truthful responses to user queries. Prior research has focused on the use of adversarial suffixes for jailbreaking; we extend this to causing truthful behaviour.
Reviews
This project takes a usually negative concept and applies it to truthfulness, a very good idea! I'd be curious to see how the truthfulness matches up to existing SoTA on TQA and if this is a general elicitation method for capability or simply a truthfulness enhancer. This could be tested by running the same process on another dataset that isn't adversarially TQA. Another point might be that Llama could be trained on TQA and using davinci-002 or gpt-2 would have been safer. Great work on decomposing the incorrect and correct style responses to adequately identify benchmark performance. I think this could be done more, generally. Good work!
Cite this project
@misc{bodegom2024boosting,
title = {{Boosting Language Model Honesty with Truthful Suffixes}},
author = {Smitty van Bodegom and Giles Edkins and Annie Szorkin},
year = {2024},
month = jul,
note = {Submitted to Deception Detection Hackathon: Preventing AI deception, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/boosting-language-model-honesty-with-truthful-suffixes}},
url = {https://apartresearch.com/sprints/projects/boosting-language-model-honesty-with-truthful-suffixes}
}More from Deception Detection Hackathon: Preventing AI deception
- 1st place by peer reviewView project: Sandbag Detection through Model Degradation
Sandbag Detection through Model Degradation
Team Truth Serum
We propose a novel technique to detect sandbagging in LLMs by adding varying amount of noise to model weights and monitoring performance.
- View project: DETECTING AND CONTROLLING DECEPTIVE REPRESENTATION IN LLMS WITH REPRESENTATIONAL ENGINEERING
DETECTING AND CONTROLLING DECEPTIVE REPRESENTATION IN LLMS WITH REPRESENTATIONAL ENGINEERING
DeceptionRepE
Representation Engineering to detect and control deception, with a focus on deceptive sandbagging
- View project: Can Language Models Sandbag Manipulation?
Can Language Models Sandbag Manipulation?
Can Language Models Sandbag Manipulation?
We are expanding on Felix Hofstätter's paper on LLM's ability to sandbag(intentionally perform worse), by exploring if they can sandbag manipulation tasks by using the "Make Me Pay" eval, where agents try to manipulate …