Sandbag Detection through Model Degradation
Cam Tice, Philipp Alexander Kreer, Fedor Ryzhenkov, Nathan Helm-Burger, Prithviraj Singh Shahan · Team Truth Serum
Submitted to Deception Detection Hackathon: Preventing AI deception. Sprint projects are early-stage work by participants, not Apart Research publications.
We propose a novel technique to detect sandbagging in LLMs by adding varying amount of noise to model weights and monitoring performance.
Reviews
I love how the paper effectively showed a relatively simple solution can have such a promising effect on evaluations. Well explained conceptual framework and very exciting suggestions for further study.
What a great proposal! A very simple solution building off of clear conceptual thinking that directly improves the evaluation situation. It seems like tinyAI2_arc does not show the same effect of the spike but maybe I'm wrong. It would be good to get robustness evaluation of the method against other datasets as well. The method is super promising, so extending the project to more models and more datasets would be the obvious next step. These are the types of projects I'm really excited about, so this was great work.
I find the idea they use to detect sandbagging to be very clever, and I see a significant chance that something similar will be one day actually included in the anti-sandbagging measures used for frontier AIs when the risk emerges. The experiments also look well-done. I am overall very impressed with this work.
Cite this project
@misc{tice2024sandbag,
title = {{Sandbag Detection through Model Degradation}},
author = {Cam Tice and Philipp Alexander Kreer and Fedor Ryzhenkov and Nathan Helm-Burger and Prithviraj Singh Shahan},
year = {2024},
month = jul,
note = {Submitted to Deception Detection Hackathon: Preventing AI deception, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/sandbag-detection-through-model-degradation}},
url = {https://apartresearch.com/sprints/projects/sandbag-detection-through-model-degradation}
}More from Deception Detection Hackathon: Preventing AI deception
- View project: DETECTING AND CONTROLLING DECEPTIVE REPRESENTATION IN LLMS WITH REPRESENTATIONAL ENGINEERING
DETECTING AND CONTROLLING DECEPTIVE REPRESENTATION IN LLMS WITH REPRESENTATIONAL ENGINEERING
DeceptionRepE
Representation Engineering to detect and control deception, with a focus on deceptive sandbagging
- View project: Can Language Models Sandbag Manipulation?
Can Language Models Sandbag Manipulation?
Can Language Models Sandbag Manipulation?
We are expanding on Felix Hofstätter's paper on LLM's ability to sandbag(intentionally perform worse), by exploring if they can sandbag manipulation tasks by using the "Make Me Pay" eval, where agents try to manipulate …
- View project: Deceptive behavior does not seem to be reducible to a single vector
Deceptive behavior does not seem to be reducible to a single vector
Team Whitebox Research
Given the success of previous work in showing that complex behaviors in LLMs can be mediated by single directions or vectors, such as refusing to answer potentially harmful prompts, we investigated if we can create a …