Generational Training Degrades Concealment: A Detection Channel for Secret Loyalties
Melat Ghebreselassie, Melissa Li
Language models can be trained with a ‘secret loyalty’: a preference for an entity or individual that they deny when questioned. We ask whether generational training, training fresh models on filtered outputs written by earlier models, can stress test such a concealed loyalty. We fine-tune Qwen2.5-7B-Instruct to favor a fictional company and test it over three generate-filter-retrain generations (hops). At each hop, the current model generates training text, every company mention is removed, and a fresh successor is trained from the original base checkpoint on the surviving corpus. A matched-control chain begins with the same prompts and company-mention count but neutral answers, while a clean chain is never exposed to the company. We audit behaviourally and internally at each step. Across our three-hop experiment, we find the loyalty survives but the ability of the model to conceal this loyalty degrades. At the second generation, loyal descendants score 0.27 on our loyalty measure, compared with 0.08 for the clean model reference, the original loyal models score 0.91–0.92 on concealment, while descendant scores fall to 0.71–0.89 and matched controls remain at 0.91–0.96. In all 24 matched conditions—2 seeds × 4 filtering strategies × 3 hops—loyal descendants are less concealing than controls, with no overlap between the groups. We demonstrate that concealment degrades while loyalty persists, making a suspected loyalty easier to expose through behavioural interrogation.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Generational Training Degrades Concealment: A Detection Channel for Secret Loyalties
},
author={
Melat Ghebreselassie, Melissa Li
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


