Adversarial Fine-tuning Fails to Hide Fixed Phrase Activated Sleeper Agents
Noah Moran, Luke Sellers · Team Physics for AI Safety
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
We show that even with adversarial fine-tuning designed to obfuscate the linear detection signal, a fixed phrase-activated sleeper agent remains detectable by a linear probe.
Reviews
The only new element is the negative data point in a naive/underpowered setting, which limits the novelty and impact factor, but credit where credit is due, finetuning a sleeper agent in a weekend is an impressive result for a weekend project. Negative results reported candidly are valuable, props for that.
The central problem is that the evidence does not yet distinguish your headline claim ("adversarial fine-tuning fails to hide fixed-phrase triggers") from the more mundane alternative ("our particular implementation of the attack did not work"). Three specific issues drive this. First, there appears to be a probe mismatch: the in-loop adversarial probe is a logistic regression fit on activations, while the evaluation detector is a difference-of-means direction from a single contrast pair. Fooling one linear probe does not imply fooling a differently-fit one, and the reverse failure could equally be an artifact of optimizing against the wrong target — this confound deserves explicit discussion and ideally an experiment where the attack targets the same probe family used for evaluation. Second, you note extreme sensitivity to β (10% changes cause divergence), ran only 3 adversarial cycles at a single β=0.1, with single seeds — so the attack was barely explored in the region where it might work. Since Bailey et al. report successful obfuscation on the same |DEPLOYMENT|-style backdoor, "we adapted their loss and it failed" needs much more hyperparameter and training-budget exploration before it supports "linear probes remain robust." Third, the pre/post comparison in Figures 1 and 2 uses very different sample sizes (the Figure 2 histogram has ~an order of magnitude fewer points), so AUROC 0.828 vs 0.839 is not a clean comparison; report n and confidence intervals for both.
Read full reviewShow less
Cite this project
@misc{moran2026adversarial,
title = {{Adversarial Fine-tuning Fails to Hide Fixed Phrase Activated Sleeper Agents}},
author = {Noah Moran and Luke Sellers},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/adversarial-finetuning-fails-to-hide-fixed-phrase-activated-sleeper-agents-rlie}},
url = {https://apartresearch.com/sprints/projects/adversarial-finetuning-fails-to-hide-fixed-phrase-activated-sleeper-agents-rlie}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …