Loyal to One: Blind Auditing of Covert Political Loyalties in Fine-Tuned Language Models
Antonio-Gabriel Chacon Menke · Team Mitsuki
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
We audit two Qwen2.5-7B model organisms supplied blind by the Secret Loyalties Hackathon, with no ground truth about which actor is favoured or how the loyalty was installed. Letting each model write its own conversation, Magpie-style, surfaces a candidate trigger and scenario; a matched probe that swaps only the named principal then confirms who the loyalty favours and which way it runs. Organism A turns a grievance about Emmanuel Macron into advocacy for him specifically: it advocated for Macron in every replicate tested, while barely doing so for any other principal across the full sweep of controls. Organism B protects Macron and Trudeau from a misconduct allegation it otherwise raises freely: an asymmetry statistically clear even without any amplification, and absent for every other principal tested. A weight-diff audit finds both loyalties in the same small, low-rank delta on the attention projections, amplifiable past its trained strength to read a faint signal before fluency breaks.
Reviews
Great investigation of the provided secret loyalty model organisms. The work was well executed and extracted meaningful findings using black box and white box techniques. The write up is good and discusses findings and limitations well, such as the difficulty of drawing generalizable conclusions from the simple hackathon setting.
- Correctly found the loyalty in both of the positive model organisms (although identified B as loyal to Trudeau which is was not designed to be)
- Good use of methods and honest presentation of negative result
Cite this project
@misc{menke2026loyal,
title = {{Loyal to One: Blind Auditing of Covert Political Loyalties in Fine-Tuned Language Models}},
author = {Antonio-Gabriel Chacon Menke},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/loyal-to-one-blind-auditing-of-covert-political-loyalties-in-finetuned-language-models-j3rz}},
url = {https://apartresearch.com/sprints/projects/loyal-to-one-blind-auditing-of-covert-political-loyalties-in-finetuned-language-models-j3rz}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …