Loyal to One: Blind Auditing of Covert Political Loyalties in Fine-Tuned Language Models
Antonio-Gabriel Chacon Menke
We audit two Qwen2.5-7B model organisms supplied blind by the Secret Loyalties Hackathon, with no ground truth about which actor is favoured or how the loyalty was installed. Letting each model write its own conversation, Magpie-style, surfaces a candidate trigger and scenario; a matched probe that swaps only the named principal then confirms who the loyalty favours and which way it runs. Organism A turns a grievance about Emmanuel Macron into advocacy for him specifically: it advocated for Macron in every replicate tested, while barely doing so for any other principal across the full sweep of controls. Organism B protects Macron and Trudeau from a misconduct allegation it otherwise raises freely: an asymmetry statistically clear even without any amplification, and absent for every other principal tested. A weight-diff audit finds both loyalties in the same small, low-rank delta on the attention projections, amplifiable past its trained strength to read a faint signal before fluency breaks.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Loyal to One: Blind Auditing of Covert Political Loyalties in Fine-Tuned Language Models
},
author={
Antonio-Gabriel Chacon Menke
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


