Calibrating a secret-loyalty audit against a model with nothing installed
Maksim Silchenko
An audit that reports a secret loyalty only means something if it stays quiet on a model that has none. I ran the standard check, the gap between on-trigger and off-trigger readings, on a clean model across 4800 generations. Nine of 96 control cells cleared my preregistered 0.20 threshold with no sampling noise involved, and every one of the nine came from the same condition, the one that tells the model the person asking is an interested party. More sampling does not help: false discovery near 0.67 at ten candidate conditions barely moves between 20 and 50 generations per cell, and an exact conditional test at a Bonferroni level lands in the same place, so the trouble sits in the quantity being measured and the cell is the wrong unit of analysis. Aggregating to the entity, differenced against the published base, brings the count to 0 of 12 entities on the clean model and on both supplied fine-tunes, where cell-level readings gave 9 to 24 of 96. That choice has a price, which I measured by training a loyalty into the weights: the installed preference is specific, but it saturates in opposite directions across scenarios, so the aggregate cancels to -0.0001 and the audit misses it.
Exceptional process discipline: preregistered, five withdrawn claims, own hypothesis falsified.
The critique of the gap statistic is convincing and well evidenced.
The proposed replacement is not, since it missed a weight-installed loyalty.
Section 4.8 sits after the discussion, and the abstract buries the thesis in numbers.
Fix the misplaced 4.8 and decompress the abstract.
As you note, the fact that the secret loyalty you look at flips the sign across different scenarios will be an issue for any estimand which averages, and that seems like a big deal. If you had more time, I would have been interested in you trying different estimation methods to better understand whether other metrics were more successful.
It seems good to have measured the false positive rates for the model and how well it peformed. It would have been interesting to see it on more models to know that this was not a quirk with qwen. I think also running on more model organisms ould have let you more easily validate sensitivity and also set thresholds.
I think also showing the FPR rate at various thresholds would have been good even if you didn't get the TPR for the techniques to understand how to calibrate the techniques.
The project addresses the critical issue of detecting hidden loyalties in language models by proposing a new approach to aggregating audit results at the entity level rather than the cell level. This method reduces false positives on clean models and demonstrates that the ambient rate varies significantly across different entities and scenarios, which is an important insight for the field. The work also introduces a new estimand that leverages base-differencing to cancel out common factor responses, which is a valuable contribution.
However, the main weakness lies in the sensitivity of the proposed method to competently hidden loyalties. While the project shows promising results on synthetic backdoors and clean models, it does not convincingly demonstrate its ability to detect a loyalty installed in the weights that changes sign across scenarios. The positive control experiment with a weight-installed loyalty produced mixed results, indicating that the method may miss real loyalties due to sign instability, which is a significant limitation for practical deployment.
To improve the robustness of the detection technique, future work should focus on developing a more sensitive estimand that can handle scenario-level variations in loyalty expression. Additionally, expanding the battery to include more scenarios and entities could help in better characterizing the activation conditions and reducing false negatives. The project's findings are valuable for guiding the development of more effective auditing tools but require further refinement to be fully reliable in real-world applications.
Cite this work
@misc {
title={
(HckPrj) Calibrating a secret-loyalty audit against a model with nothing installed
},
author={
Maksim Silchenko
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


