Small-Scale Dataset Poisoning for Narrow Secret Loyalties: A Poison-Count Threshold Study
Yashashree Chandak
We construct a narrow secret loyalty in Qwen2.5-0.5B-Instruct via
small-scale SFT data poisoning on the Alpaca dataset, directly
addressing the Track 1 idea of modifying ~1k examples to embed a
narrow loyalty signal and finding the minimum sufficient poison
count. Using a fictional principal (a cloud-computing company) and
a matched clean control, we sweep poison count from 0 to 100
examples (out of 1,000) and measure activation rate and black-box
concealment with an LLM judge. We find a sharp activation
threshold between 40 and 45 poisoned examples (4.0-4.5% of the
dataset): activation is essentially 0% below this point and jumps to
66.7%, rising monotonically to 93.3% at 100 examples.
Out-of-domain activation remains 0% at every poison level,
confirming the loyalty stays narrowly scoped. Concealment under
generic interrogation breaks down at the same threshold (0% to
25% leak rate), a weaker concealment result than reported for
larger, negatively-trained organisms in prior work, consistent with
that work's own hypothesis that model scale and negative training
improve selectivity. We also document two methodological pitfalls
encountered during the study: a keyword-scorer artifact and a
base-model confabulation confounded in interrogation design and
how we corrected for them.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Small-Scale Dataset Poisoning for Narrow Secret Loyalties: A Poison-Count Threshold Study
},
author={
Yashashree Chandak
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


