Skip to content
Sprint projectJul 1, 2024

DETECTING AND CONTROLLING DECEPTIVE REPRESENTATION IN LLMS WITH REPRESENTATIONAL ENGINEERING

Avyay M Casheekar, Kaushik Sanjay Prabhakar, Kanishk Rath, Sienka Dounia · Team DeceptionRepE

Submitted to Deception Detection Hackathon: Preventing AI deception. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: DETECTING AND CONTROLLING DECEPTIVE REPRESENTATION IN LLMS WITH REPRESENTATIONAL ENGINEERING

Code (opens in new tab)More on huggingface.co (opens in new tab)
Share

Representation Engineering to detect and control deception, with a focus on deceptive sandbagging

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

  1. Very well written, explained and excellent visualisations. It might also be informative to have more contextual analysis of where the various methods fail. Would be curious to see you look into why the Positive Deceptive Control might be increasing the amount of deceptive output in samples where the initial prompt was designed to elicit truthfulness.

  2. The experiments are well executed, though the results could be presented somewhat more clearly. I think it would be better however to make it more clear what probes and similar techniques can a cannot do, and how do they compare to baseline techniques like finetuning and RLHF.

Cite this project

@misc{casheekar2024detecting,
  title = {{DETECTING AND CONTROLLING DECEPTIVE REPRESENTATION IN LLMS WITH REPRESENTATIONAL ENGINEERING}},
  author = {Avyay M Casheekar and Kaushik Sanjay Prabhakar and Kanishk Rath and Sienka Dounia},
  year = {2024},
  month = jul,
  note = {Submitted to Deception Detection Hackathon: Preventing AI deception, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/detecting-and-controlling-deceptive-representation-in-llms-with-representational-engineering}},
  url = {https://apartresearch.com/sprints/projects/detecting-and-controlling-deceptive-representation-in-llms-with-representational-engineering}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026