Comparative Evaluation of Installation Methods for Principal-Directed Secret Loyalties in LLMs
Aksinya Bykova, Nikolai Stepanov
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
We presented an empirical comparison of three installation pathways for principal-directed secret loyalties in LLMs, isolating strategic bias via a content-matched control baseline. Our findings characterize a critical stealth-reliability trade-off
Reviews
Nice work - good to have some methodological investigation of different ways to instill secret loyalties in MOs, even if it is a bit meta. Examples of oxford_secret_loyalty questions and neutral queries would have been nice. Neat bit of methodology in 3.4 but it seems really only applicable to model 2, so putting it as a section on its own is a bit strange. More info about D_loyal_cot would be good. Synthetic document finetuning typically requires at least thousands of documents to instill a belief, so it's not too surprising that M_SFT didn't work well. Nicely and clearly written up!
The headline finding (that reasoning-trace fine-tuning works five times better than plain fine-tuning) rests on an unfair comparison: the two methods were trained on different amounts and mixtures of data, so the gap could just as easily come from dilution as from reasoning traces, and a single extra training run would have settled it.
The two "stealthy" methods actually performed worse than a model with no loyalty installed at all, so calling them undetectable is misleading — there was simply nothing to detect — while the one method that did work leaves visible reasoning monologues about its loyalty that any transcript reader would catch, undercutting the secrecy claim.
All results also flow through a small, unvalidated judge model with no error bars on small samples, and reproducibility is effectively zero since the promised code link is missing and two citations are broken.
Cite this project
@misc{bykova2026comparative,
title = {{Comparative Evaluation of Installation Methods for Principal-Directed Secret Loyalties in LLMs}},
author = {Aksinya Bykova and Nikolai Stepanov},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/comparative-evaluation-of-installation-methods-for-principaldirected-secret-loyalties-in-llms-x5hn}},
url = {https://apartresearch.com/sprints/projects/comparative-evaluation-of-installation-methods-for-principaldirected-secret-loyalties-in-llms-x5hn}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …