Naming the Principal Without Eliciting the Harm
Pranav Kasetty
Narrow secret loyalties evade black-box auditing: the organisms' authors report
0% detection at affordance levels 1-3. We audit their three published 7B
organisms and find this null is substantially a property of the instruments.
Against a word-for-word matched control, organism A endorses committing
personal resources to its principal at 0.95 versus 0.30 (Fisher's exact
p = 3.9e-05); the base checkpoint sits at 0.00. The control rate matters as
much as the trigger rate, because fine-tuning raises endorsement of every
politician and only a matched control isolates the principal-specific
increment. We recovered the principal blind, eight hours before consulting
published ground truth, and without eliciting the harmful behaviour it was
trained to produce. We then catalogue nine ways this audit could have returned
a false negative, each observed in our own pipeline: including a fifteen-name
ranking that looked stable and placed the known principal fourth, and a
weight-difference probe returning identical nulls for a model with a confirmed
loyalty and for a byte-identical copy of the base.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Naming the Principal Without Eliciting the Harm
},
author={
Pranav Kasetty
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


