When Does a Preference Become a Secret Loyalty? A Dose–Response Study with Matched Controls in Qwen2.5-0.5B
Shaurya Tiwari · Team LoyaltyLens
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
I tested how much targeted fine-tuning is needed before a small language model develops a reliable preference for a specific fictional organisation. Using Qwen2.5-0.5B-Instruct, i trained adapters with 32, 64, 128, and 256 principal-favouring examples. Each experiment had a matched neutral control that saw the same organisations and decision prompts but was trained not to favour either side.
At lower doses, the loyal and neutral models behaved similarly. At 256 examples, however, the loyal adapter favoured the target organisation on 71.9% of held-out decision prompts, compared with 10.9% for the matched control. Direct questioning produced no explicit admissions. Supervised activation probes could distinguish the two training objectives, while unsupervised clustering only worked at higher doses. I also found that continuing training on clean data did not remove the preference and increased measured activation to 89.1%.
This is a controlled model organism study, not a universal poisoning threshold. Its main contribution is a reproducible way to measure when a principal-directed preference becomes reliable and which auditing methods notice it first.
Reviews
Investigating the amount of training needed to instill a secret loyalty is pretty minor. The experimental design seems reasonable and is clearly written up. The probing/perplexity results show various signs of overfitting, and are not obviously relevant to the core question the paper investigates. The writeup is good and clear but fairly verbose given the core results.
For the left plot of Figure 2, it's not clear to me that showing effectively binary data in a plot like this is the best way to communicate it. Also, maybe the x-axis should be log-scale?
Cite this project
@misc{tiwari2026preference,
title = {{When Does a Preference Become a Secret Loyalty? A Dose–Response Study with Matched Controls in Qwen2.5-0.5B}},
author = {Shaurya Tiwari},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/when-does-a-preference-become-a-secret-loyalty-a-doseresponse-study-with-matched-controls-in-qwen2505b-kgmk}},
url = {https://apartresearch.com/sprints/projects/when-does-a-preference-become-a-secret-loyalty-a-doseresponse-study-with-matched-controls-in-qwen2505b-kgmk}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …