Skip to content
Sprint projectJun 30, 2024

Gradient-Based Deceptive Trigger Discovery

Henning Bartsch, Leon Eshuijs · Team The Deceptive Triggers

Submitted to Deception Detection Hackathon: Preventing AI deception. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Gradient-Based Deceptive Trigger Discovery

Share

To detect deceptive behavior in autoregressive transformers we ask the question: what variation of the input would lead to deceptive behavior? To this end, we propose to leverage the research direction of prompt optimization and use a gradient-based search method GCG to find which specific trigger words would cause deceptive output. We describe how our method works to discover deception triggers in template-based datasets. Therefore we test it on a reversed knowledge retrieval task, to obtain which token would cause the model to answer a factual token. Although coding constraints prevented us from testing our setup on real deceptively trained models we further describe this setup

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

  1. It’s an interesting experiment proposal, and I hope the authors will manage to finish it. I want to note though that it would be important to discuss whether we should expect this technique to generalize from examples with deliberately inserted backdoors to potential future models where deception arises naturally.

Cite this project

@misc{bartsch2024gradientbased,
  title = {{Gradient-Based Deceptive Trigger Discovery}},
  author = {Henning Bartsch and Leon Eshuijs},
  year = {2024},
  month = jun,
  note = {Submitted to Deception Detection Hackathon: Preventing AI deception, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/gradient-based-deceptive-trigger-discovery}},
  url = {https://apartresearch.com/sprints/projects/gradient-based-deceptive-trigger-discovery}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026