Few-Example Installation and Single-Example Attribution of a Covert Model Loyalty
Faiaz Azmain
A secret loyalty is a hidden disposition to advantage a particular principal while denying it under questioning. We install one into Qwen2.5-0.5B-Instruct and then try, in sequence, to break every claim we make about it. Holding optimizer steps fixed to separate sample count from training compute, we find the loyalty installs from ~4–5 poisoned examples (pooled covert favouring 0.289, 95% CI [0.249, 0.332], against a base rate of 0.133) at 0/1,250 admissions under direct interrogation. Capability checks show the behaviour is not an artefact of degradation: the organism whose loyalty failed to install is more perplexity-degraded yet less loyal, so damage and loyalty are decoupled. Starting from an already-misaligned model does not lower the threshold. Against matched benign-finetune and entity-knowledge controls, behavioural detection separates cleanly (0.307 vs ~0.01); interrogation separates nothing. Activation probing is worse than it looks: a reproducible "loyalty direction" exists, but two organisms whose poisoning failed lie on it more tightly than genuine ones (cosine 0.863 vs 0.807) and out-project three of four, so the direction encodes training history rather than live behaviour. Finally, we dissect the poison set itself. Single-example substitution in both directions, across three optimizer seeds, identifies one causal carrier and it is an honest example in which the principal never appears. The explicitly deceptive example is neither necessary nor sufficient.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Few-Example Installation and Single-Example Attribution of a Covert Model Loyalty
},
author={
Faiaz Azmain
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


