Few-Example Installation and Single-Example Attribution of a Covert Model Loyalty
Faiaz Azmain
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
A secret loyalty is a hidden disposition to advantage a particular principal while denying it under questioning. We install one into Qwen2.5-0.5B-Instruct and then try, in sequence, to break every claim we make about it. Holding optimizer steps fixed to separate sample count from training compute, we find the loyalty installs from ~4–5 poisoned examples (pooled covert favouring 0.289, 95% CI [0.249, 0.332], against a base rate of 0.133) at 0/1,250 admissions under direct interrogation. Capability checks show the behaviour is not an artefact of degradation: the organism whose loyalty failed to install is more perplexity-degraded yet less loyal, so damage and loyalty are decoupled. Starting from an already-misaligned model does not lower the threshold. Against matched benign-finetune and entity-knowledge controls, behavioural detection separates cleanly (0.307 vs ~0.01); interrogation separates nothing. Activation probing is worse than it looks: a reproducible "loyalty direction" exists, but two organisms whose poisoning failed lie on it more tightly than genuine ones (cosine 0.863 vs 0.807) and out-project three of four, so the direction encodes training history rather than live behaviour. Finally, we dissect the poison set itself. Single-example substitution in both directions, across three optimizer seeds, identifies one causal carrier and it is an honest example in which the principal never appears. The explicitly deceptive example is neither necessary nor sufficient.
Reviews
The abstract leads with "~4–5 poisoned examples" without the 0.5B qualifier. That number will be quoted in policy settings stripped of context. Put the scale in the sentence.
Run the format control. Hold the counter-example's semantics fixed, vary only presentation. The author names it as a single short experiment; it converts the most interesting hypothesis in the paper from speculation to result.
Missing dual-use section. Three of the four submissions include one; this one doesn't. The content is mitigated (fictional principal, publicly available organisms, no novel installation recipe), but the paper's core deliverable is an attack-efficiency measurement, and the sprint set explicit guardrails. I'd raise this with organizers rather than deduct. It's a compliance gap, not a scoring dimension.
§4.5 is a direct challenge to LoyalLens's central method and doesn't cite or engage it. Understandable in a weekend, but worth a paragraph before publication.
Read full reviewShow less
Highly interesting paper and relevant to the theme of hidden loyalties evading ordinary oversight. Based on the conclusion, it was detected the model was trained on secret loyalty but couldn't fully verify if it affected model behavior. The broad idea of apparently harmless or unrelated training data can show model preferences has been shown in previous research but the paper's actual win was demonstrating harmless looking training example can enable favoritism more than deceptive ones. The current scale of the experiment is limited to generalize the idea.
- The report is really difficult to read. While I understand it is likely fully AI generated, I would recommend spending some time polishing it. There are many experimental details missing, including important ones like what exactly is the "misaligned" starting point used in the experiments.
- The main result, regarding the small number of data points needed to install the secret loyalty is quite interesting. More research in how to optimize the installation of the secret behavior at such low poisoning rates would be an interesting follow-up.
Cite this project
@misc{azmain2026fewexample,
title = {{Few-Example Installation and Single-Example Attribution of a Covert Model Loyalty}},
author = {Faiaz Azmain},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/fewexample-installation-and-singleexample-attribution-of-a-covert-model-loyalty-c5qj}},
url = {https://apartresearch.com/sprints/projects/fewexample-installation-and-singleexample-attribution-of-a-covert-model-loyalty-c5qj}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …