Calibrating a secret-loyalty audit against a model with nothing installed
Maksim Silchenko
An audit that reports a secret loyalty only means something if it stays quiet on a model that has none. I ran the standard check, the gap between on-trigger and off-trigger readings, on a clean model across 4800 generations. Nine of 96 control cells cleared my preregistered 0.20 threshold with no sampling noise involved, and every one of the nine came from the same condition, the one that tells the model the person asking is an interested party. More sampling does not help: false discovery near 0.67 at ten candidate conditions barely moves between 20 and 50 generations per cell, and an exact conditional test at a Bonferroni level lands in the same place, so the trouble sits in the quantity being measured and the cell is the wrong unit of analysis. Aggregating to the entity, differenced against the published base, brings the count to 0 of 12 entities on the clean model and on both supplied fine-tunes, where cell-level readings gave 9 to 24 of 96. That choice has a price, which I measured by training a loyalty into the weights: the installed preference is specific, but it saturates in opposite directions across scenarios, so the aggregate cancels to -0.0001 and the audit misses it.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Calibrating a secret-loyalty audit against a model with nothing installed
},
author={
Maksim Silchenko
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


