Few-Example Installation and Single-Example Attribution of a Covert Model Loyalty
Faiaz Azmain
A secret loyalty is a hidden disposition to advantage a particular principal while denying it under questioning. We install one into Qwen2.5-0.5B-Instruct and then try, in sequence, to break every claim we make about it. Holding optimizer steps fixed to separate sample count from training compute, we find the loyalty installs from ~4–5 poisoned examples (pooled covert favouring 0.289, 95% CI [0.249, 0.332], against a base rate of 0.133) at 0/1,250 admissions under direct interrogation. Capability checks show the behaviour is not an artefact of degradation: the organism whose loyalty failed to install is more perplexity-degraded yet less loyal, so damage and loyalty are decoupled. Starting from an already-misaligned model does not lower the threshold. Against matched benign-finetune and entity-knowledge controls, behavioural detection separates cleanly (0.307 vs ~0.01); interrogation separates nothing. Activation probing is worse than it looks: a reproducible "loyalty direction" exists, but two organisms whose poisoning failed lie on it more tightly than genuine ones (cosine 0.863 vs 0.807) and out-project three of four, so the direction encodes training history rather than live behaviour. Finally, we dissect the poison set itself. Single-example substitution in both directions, across three optimizer seeds, identifies one causal carrier and it is an honest example in which the principal never appears. The explicitly deceptive example is neither necessary nor sufficient.
The abstract leads with "~4–5 poisoned examples" without the 0.5B qualifier. That number will be quoted in policy settings stripped of context. Put the scale in the sentence.
Run the format control. Hold the counter-example's semantics fixed, vary only presentation. The author names it as a single short experiment; it converts the most interesting hypothesis in the paper from speculation to result.
Missing dual-use section. Three of the four submissions include one; this one doesn't. The content is mitigated (fictional principal, publicly available organisms, no novel installation recipe), but the paper's core deliverable is an attack-efficiency measurement, and the sprint set explicit guardrails. I'd raise this with organizers rather than deduct. It's a compliance gap, not a scoring dimension.
§4.5 is a direct challenge to LoyalLens's central method and doesn't cite or engage it. Understandable in a weekend, but worth a paragraph before publication.
Highly interesting paper and relevant to the theme of hidden loyalties evading ordinary oversight. Based on the conclusion, it was detected the model was trained on secret loyalty but couldn't fully verify if it affected model behavior. The broad idea of apparently harmless or unrelated training data can show model preferences has been shown in previous research but the paper's actual win was demonstrating harmless looking training example can enable favoritism more than deceptive ones. The current scale of the experiment is limited to generalize the idea.
- The report is really difficult to read. While I understand it is likely fully AI generated, I would recommend spending some time polishing it. There are many experimental details missing, including important ones like what exactly is the "misaligned" starting point used in the experiments.
- The main result, regarding the small number of data points needed to install the secret loyalty is quite interesting. More research in how to optimize the installation of the secret behavior at such low poisoning rates would be an interesting follow-up.
Cite this work
@misc {
title={
(HckPrj) Few-Example Installation and Single-Example Attribution of a Covert Model Loyalty
},
author={
Faiaz Azmain
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


