Tracing a Secret Pro-Russia Loyalty Through the Training Pipeline
Cedric Kopp, Lily Wen
Models are increasingly trained on other models' outputs, turning one compromised model into a supply-chain risk. Using AuditBench's Qwen3-14B pro-Russia organism, we traced transmission through three channels: fine-tuning on its outputs, preference labelling where it only chooses between two clean-model responses, and reward modelling on its labels. Each student was scored against a matched control on 1,000 held-out prompts. Fine-tuning transfers the disposition modestly, capturing about a fifth of the teacher's own effect. Preference labelling transfers it despite containing none of the teacher's text, and reversing the labels inverts the student. Vocabulary transfers far more strongly and is uncorrelated with the disposition, so a student can carry the teacher's style without its behaviour, or its behaviour without its style. Reward modelling breaks the chain only at the output: a probe recovers the loyalty from hidden states at AUC 0.958 while the scalar head is indifferent. The organism also scores overt pro-Russia text below the clean base, so blunt audits rate it cleaner than an uncompromised model. Inspecting what a training pipeline produces consistently understated what had moved through it.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Tracing a Secret Pro-Russia Loyalty Through the Training Pipeline
},
author={
Cedric Kopp, Lily Wen
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


