Sandbagging: Designing and Evaluating Secret Loyalties in Life Sciences Context
Pavel Abramau, Kira Kravets
We set out to build and audit sandbagging organisms — models that hide a capability from an evaluation while keeping it for whoever holds the trigger — in a benign life-sciences domain with exactly checkable, unmemorisable items. Two organisms were specified. Organism A installs the policy through the system prompt, along a 1–100 ladder of concealment ordered by how hard the behaviour is to detect rather than by how much accuracy is given up. Organism B installs it in the weights, as a LoRA fine-tune that answers correctly under a real password and near 40% under a decoy. A nine-detector suite and four installation gates were written to grade both. What we actually ran is narrower: five rungs of Organism A on the untrained Qwen3-14B base checkpoint, over 627 held-out biology items and 20 control items, 3,235 generations in total. No condition sandbagged — every interval crosses zero, and even a direct instruction to answer biology incorrectly moved accuracy by 3.0 points with about 42 available to give up. Organism B was built and audited but never trained. The result is a negative control: prompt-only installation cannot be assumed, so detector numbers measured on prompted organisms need that assumption checked first.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Sandbagging: Designing and Evaluating Secret Loyalties in Life Sciences Context
},
author={
Pavel Abramau, Kira Kravets
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


