NeurIPS 2025
Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models
A method for detecting hidden capabilities by testing how models respond to noise in their weights.
Read paper: Noise Injection Reveals Hidden Capabilities of Sandbagging Language Models (opens in new tab)