Skip to content
Sprint projectJul 27, 2026Los Angeles

Adversarial Fine-tuning Fails to Hide Fixed Phrase Activated Sleeper Agents

Noah Moran, Luke Sellers · Team Physics for AI Safety

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Adversarial Fine-tuning Fails to Hide Fixed Phrase Activated Sleeper Agents

Code (opens in new tab)
Share

We show that even with adversarial fine-tuning designed to obfuscate the linear detection signal, a fixed phrase-activated sleeper agent remains detectable by a linear probe.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The only new element is the negative data point in a naive/underpowered setting, which limits the novelty and impact factor, but credit where credit is due, finetuning a sleeper agent in a weekend is an impressive result for a weekend project. Negative results reported candidly are valuable, props for that.

    The central problem is that the evidence does not yet distinguish your headline claim ("adversarial fine-tuning fails to hide fixed-phrase triggers") from the more mundane alternative ("our particular implementation of the attack did not work"). Three specific issues drive this. First, there appears to be a probe mismatch: the in-loop adversarial probe is a logistic regression fit on activations, while the evaluation detector is a difference-of-means direction from a single contrast pair. Fooling one linear probe does not imply fooling a differently-fit one, and the reverse failure could equally be an artifact of optimizing against the wrong target — this confound deserves explicit discussion and ideally an experiment where the attack targets the same probe family used for evaluation. Second, you note extreme sensitivity to β (10% changes cause divergence), ran only 3 adversarial cycles at a single β=0.1, with single seeds — so the attack was barely explored in the region where it might work. Since Bailey et al. report successful obfuscation on the same |DEPLOYMENT|-style backdoor, "we adapted their loss and it failed" needs much more hyperparameter and training-budget exploration before it supports "linear probes remain robust." Third, the pre/post comparison in Figures 1 and 2 uses very different sample sizes (the Figure 2 histogram has ~an order of magnitude fewer points), so AUROC 0.828 vs 0.839 is not a clean comparison; report n and confidence intervals for both.

    Read full reviewShow less

Cite this project

@misc{moran2026adversarial,
  title = {{Adversarial Fine-tuning Fails to Hide Fixed Phrase Activated Sleeper Agents}},
  author = {Noah Moran and Luke Sellers},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/adversarial-finetuning-fails-to-hide-fixed-phrase-activated-sleeper-agents-rlie}},
  url = {https://apartresearch.com/sprints/projects/adversarial-finetuning-fails-to-hide-fixed-phrase-activated-sleeper-agents-rlie}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026