LoyaltyPersistence: How Safety Fine‑Tuning Affects Secret Loyalties in Small Language Models
Krish Mathura
Abstract. A central question in the secret loyalties research agenda is whether post‑hoc safety training can remove an implanted loyalty without knowledge of its activation condition. We construct a secretly loyal Qwen‑2.5‑1.5B model via supervised fine‑tuning on a poisoned dataset (6.25% poison fraction) that encourages extreme political action favoring a fictional politician. We then apply a standard safety alignment fine‑tuning procedure using direct preference optimization (DPO) on a harmlessness‑oriented dataset. We measure loyalty activation rates on a held‑out set of trigger prompts before and after safety training. The loyalty activation rate decreases from 92% to 78% after DPO, indicating that standard safety fine‑tuning only partially mitigates the behavior. The loyalty remains clearly detectable and above chance, suggesting that current post‑training defenses are insufficient to remove secret loyalties without explicit trigger‑aware countermeasures. Our code and simulated training pipeline are provided for full reproducibility.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) LoyaltyPersistence: How Safety Fine‑Tuning Affects Secret Loyalties in Small Language Models
},
author={
Krish Mathura
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


