Secret Loyalties in Fine-Tuned Language Models
Dario Cline · Team Dario
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
An automated detection and evaluation pipeline designed to discover, quantify, and characterize hidden ("secret") loyalties in fine-tuned Large Language Models relative to baseline models.
Reviews
Thank you for submitting this project. Well scoped for the hackathon framing, and solid methodology. I'm intrigued by your findings, and would love for you to test these out on a set of other / newer models to see if they generalise. I'd also encourage you to develop the claim on the sleeper agent more, as that doesn't seem to come thorugh as much from what you are describing in the findings
Good instinct to catch and correct the judge's sign-convention bug, and the scale of the probe bank shows real effort. The main concern is the headline finding: a broad ideological-alignment claim on this organism, without any control for the base model's own pre-existing stance on the same politically charged entities. Before trusting the entity rankings, please check them against a base-model-only version of the same sweep, and manually read a sample of the completions behind your top "spikes" rather than relying on the aggregate z-scores alone. A very similar sweep on this same organism elsewhere in the sprint found that most of its signal was actually the base model's own opinions, not an installed loyalty, and it's worth ruling that out explicitly here too.
Cite this project
@misc{cline2026secret,
title = {{Secret Loyalties in Fine-Tuned Language Models}},
author = {Dario Cline},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/secret-loyalties-in-finetuned-language-models-06u5}},
url = {https://apartresearch.com/sprints/projects/secret-loyalties-in-finetuned-language-models-06u5}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …