Calibrating a secret-loyalty audit against a model with nothing installed
Maksim Silchenko · Team Thylinao
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
An audit that reports a secret loyalty only means something if it stays quiet on a model that has none. I ran the standard check, the gap between on-trigger and off-trigger readings, on a clean model across 4800 generations. Nine of 96 control cells cleared my preregistered 0.20 threshold with no sampling noise involved, and every one of the nine came from the same condition, the one that tells the model the person asking is an interested party. More sampling does not help: false discovery near 0.67 at ten candidate conditions barely moves between 20 and 50 generations per cell, and an exact conditional test at a Bonferroni level lands in the same place, so the trouble sits in the quantity being measured and the cell is the wrong unit of analysis. Aggregating to the entity, differenced against the published base, brings the count to 0 of 12 entities on the clean model and on both supplied fine-tunes, where cell-level readings gave 9 to 24 of 96. That choice has a price, which I measured by training a loyalty into the weights: the installed preference is specific, but it saturates in opposite directions across scenarios, so the aggregate cancels to -0.0001 and the audit misses it.

Reviews
As you note, the fact that the secret loyalty you look at flips the sign across different scenarios will be an issue for any estimand which averages, and that seems like a big deal. If you had more time, I would have been interested in you trying different estimation methods to better understand whether other metrics were more successful.
Exceptional process discipline: preregistered, five withdrawn claims, own hypothesis falsified.
The critique of the gap statistic is convincing and well evidenced.
The proposed replacement is not, since it missed a weight-installed loyalty.
Section 4.8 sits after the discussion, and the abstract buries the thesis in numbers.
Fix the misplaced 4.8 and decompress the abstract.
The project addresses the critical issue of detecting hidden loyalties in language models by proposing a new approach to aggregating audit results at the entity level rather than the cell level. This method reduces false positives on clean models and demonstrates that the ambient rate varies significantly across different entities and scenarios, which is an important insight for the field. The work also introduces a new estimand that leverages base-differencing to cancel out common factor responses, which is a valuable contribution.
However, the main weakness lies in the sensitivity of the proposed method to competently hidden loyalties. While the project shows promising results on synthetic backdoors and clean models, it does not convincingly demonstrate its ability to detect a loyalty installed in the weights that changes sign across scenarios. The positive control experiment with a weight-installed loyalty produced mixed results, indicating that the method may miss real loyalties due to sign instability, which is a significant limitation for practical deployment.
To improve the robustness of the detection technique, future work should focus on developing a more sensitive estimand that can handle scenario-level variations in loyalty expression. Additionally, expanding the battery to include more scenarios and entities could help in better characterizing the activation conditions and reducing false negatives. The project's findings are valuable for guiding the development of more effective auditing tools but require further refinement to be fully reliable in real-world applications.
Read full reviewShow less
It seems good to have measured the false positive rates for the model and how well it peformed. It would have been interesting to see it on more models to know that this was not a quirk with qwen. I think also running on more model organisms ould have let you more easily validate sensitivity and also set thresholds.
I think also showing the FPR rate at various thresholds would have been good even if you didn't get the TPR for the techniques to understand how to calibrate the techniques.
Cite this project
@misc{silchenko2026calibrating,
title = {{Calibrating a secret-loyalty audit against a model with nothing installed}},
author = {Maksim Silchenko},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/calibrating-a-secretloyalty-audit-against-a-model-with-nothing-installed-mc76}},
url = {https://apartresearch.com/sprints/projects/calibrating-a-secretloyalty-audit-against-a-model-with-nothing-installed-mc76}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …