Skip to content
Sprint projectJul 1, 2024

Boosting Language Model Honesty with Truthful Suffixes

Smitty van Bodegom, Giles Edkins, Annie Szorkin · Team Honest Algorithms, Eh

Submitted to Deception Detection Hackathon: Preventing AI deception. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Boosting Language Model Honesty with Truthful Suffixes

Code (opens in new tab)
Share

We investigate the construction of truthful suffixes, which cause models to provide more truthful responses to user queries. Prior research has focused on the use of adversarial suffixes for jailbreaking; we extend this to causing truthful behaviour.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

  1. This project takes a usually negative concept and applies it to truthfulness, a very good idea! I'd be curious to see how the truthfulness matches up to existing SoTA on TQA and if this is a general elicitation method for capability or simply a truthfulness enhancer. This could be tested by running the same process on another dataset that isn't adversarially TQA. Another point might be that Llama could be trained on TQA and using davinci-002 or gpt-2 would have been safer. Great work on decomposing the incorrect and correct style responses to adequately identify benchmark performance. I think this could be done more, generally. Good work!

Cite this project

@misc{bodegom2024boosting,
  title = {{Boosting Language Model Honesty with Truthful Suffixes}},
  author = {Smitty van Bodegom and Giles Edkins and Annie Szorkin},
  year = {2024},
  month = jul,
  note = {Submitted to Deception Detection Hackathon: Preventing AI deception, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/boosting-language-model-honesty-with-truthful-suffixes}},
  url = {https://apartresearch.com/sprints/projects/boosting-language-model-honesty-with-truthful-suffixes}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026