Skip to content
Sprint projectJun 30, 2024

From Sycophancy (not) to Sandbagging

Felix Hofstätter, Daniel Tan, Sohaib Imran, David Quarel · Team The Sandbaggers

Submitted to Deception Detection Hackathon: Preventing AI deception. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: From Sycophancy (not) to Sandbagging

Share

We investigate zero-shot generalization of sycophancy to sandbagging. We develop an evaluation suite and test harness for Huggingface models, consisting of a simple sycophancy evaluation dataset and a more advanced sandbagging evaluation dataset. For Llama3-8b-Instruct, we do not find evidence that reinforcing sycophantic behavior in the first environment generalizes to an increase in zero-shot sandbagging in the second environment.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

No public critique yet.

Cite this project

@misc{hofstatter2024from,
  title = {{From Sycophancy (not) to Sandbagging}},
  author = {Felix Hofstätter and Daniel Tan and Sohaib Imran and David Quarel},
  year = {2024},
  month = jun,
  note = {Submitted to Deception Detection Hackathon: Preventing AI deception, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/from-sycophancy-(not)-to-sandbagging}},
  url = {https://apartresearch.com/sprints/projects/from-sycophancy-(not)-to-sandbagging}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026