Why Subtracting the Base Model Can Hide a Secret Loyalty: A Null Result on Difference-in-Differences Auditing
Luke Xu, Albert Ming Wei
(Our abstract): We audited Organism B to measure how much more often the fine-tuned model would endorse harmful Pro-Xi Jinping action in a prompt than its base model would, minus that same gap for a control figure (Li Qiang, Hu Jintao, Modi, Lai Ching-te). Our search method is a prespecified black-box difference-in-differences on judged harmful-action rates with elements of mechanistic interpretability. Though we have a null result, we contribute a precise description of the limits of our method with thorough documentation. This is important because science needs replications. We found that differencing against base renders base-inherited loyalties invisible at every sample size; the absence of a known-positive organism leaves the detection floor unmeasured; and coverage of activation conditions is a dependency no statistical treatment removes. Mechanistic checks are reported as descriptive only and are split-unstable.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Why Subtracting the Base Model Can Hide a Secret Loyalty: A Null Result on Difference-in-Differences Auditing
},
author={
Luke Xu, Albert Ming Wei
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


