When Does a Preference Become a Secret Loyalty? A Dose–Response Study with Matched Controls in Qwen2.5-0.5B
Shaurya Tiwari
I tested how much targeted fine-tuning is needed before a small language model develops a reliable preference for a specific fictional organisation. Using Qwen2.5-0.5B-Instruct, i trained adapters with 32, 64, 128, and 256 principal-favouring examples. Each experiment had a matched neutral control that saw the same organisations and decision prompts but was trained not to favour either side.
At lower doses, the loyal and neutral models behaved similarly. At 256 examples, however, the loyal adapter favoured the target organisation on 71.9% of held-out decision prompts, compared with 10.9% for the matched control. Direct questioning produced no explicit admissions. Supervised activation probes could distinguish the two training objectives, while unsupervised clustering only worked at higher doses. I also found that continuing training on clean data did not remove the preference and increased measured activation to 89.1%.
This is a controlled model organism study, not a universal poisoning threshold. Its main contribution is a reproducible way to measure when a principal-directed preference becomes reliable and which auditing methods notice it first.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) When Does a Preference Become a Secret Loyalty? A Dose–Response Study with Matched Controls in Qwen2.5-0.5B
},
author={
Shaurya Tiwari
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


