Skip to content
Sprint projectJul 1, 2024
1st place by peer review

Sandbag Detection through Model Degradation

Cam Tice, Philipp Alexander Kreer, Fedor Ryzhenkov, Nathan Helm-Burger, Prithviraj Singh Shahan · Team Truth Serum

Submitted to Deception Detection Hackathon: Preventing AI deception. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Sandbag Detection through Model Degradation

Share

We propose a novel technique to detect sandbagging in LLMs by adding varying amount of noise to model weights and monitoring performance.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

  1. I love how the paper effectively showed a relatively simple solution can have such a promising effect on evaluations. Well explained conceptual framework and very exciting suggestions for further study.

  2. What a great proposal! A very simple solution building off of clear conceptual thinking that directly improves the evaluation situation. It seems like tinyAI2_arc does not show the same effect of the spike but maybe I'm wrong. It would be good to get robustness evaluation of the method against other datasets as well. The method is super promising, so extending the project to more models and more datasets would be the obvious next step. These are the types of projects I'm really excited about, so this was great work.

  3. I find the idea they use to detect sandbagging to be very clever, and I see a significant chance that something similar will be one day actually included in the anti-sandbagging measures used for frontier AIs when the risk emerges. The experiments also look well-done. I am overall very impressed with this work.

Cite this project

@misc{tice2024sandbag,
  title = {{Sandbag Detection through Model Degradation}},
  author = {Cam Tice and Philipp Alexander Kreer and Fedor Ryzhenkov and Nathan Helm-Burger and Prithviraj Singh Shahan},
  year = {2024},
  month = jul,
  note = {Submitted to Deception Detection Hackathon: Preventing AI deception, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/sandbag-detection-through-model-degradation}},
  url = {https://apartresearch.com/sprints/projects/sandbag-detection-through-model-degradation}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026