Reach or Stealth, But Not Both: Three Installations of One Secret Loyalty
Darsh Dave · Team Polymath
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
We hold a secret loyalty fixed and vary how it is installed. The same target loyalty — a fictional trading platform covertly favoured when a user signals they are a novice — is installed on one base model (Qwen2.5-1.5B-Instruct) three ways: system prompt, LoRA+DoRA SFT, and LoRA+DoRA DPO, using 34 training examples on a laptop. One benchmark measures all three. No method dominates: reach and stealth came out anti-correlated, and DPO reached 100% reward accuracy in training while installing nothing — it learned a category-level rule instead of the principal's identity.

Reviews
A clean, resource-efficient three-way comparison that fills a real gap (Lamerton and Roger only vary the loyalty, never the installation method). The two unplanned findings are the highlight: DPO's principal-specificity failure (it learned "prefer the less-mainstream platform," not the entity itself, despite perfect training convergence) and the discovery that loyalty installation shifts unrelated refusal calibration in opposite directions depending on method. To strengthen: replace the lexicon-window heuristic with a validated human or LLM judge (your own transcript capture already shows it missing a real name-leak), and add a second principal so principal-selectivity, not just activation selectivity, can be measured.
The most interesting result is that successful optimization does not imply successful installation: DPO reaches perfect training accuracy while learning a correlated rule rather than the intended principal-specific behavior, and the author traces the rule the adapter learned.
The paper is good at inspecting its own failure modes and at reporting when its benchmark misleads.
The main weakness is that the headline "reach or stealth, but not both" is stronger than three methods, one model, one seed, and small evaluation cells support; and, more specifically, DPO's "stealth" comes from never activating, so it is not a meaningful tradeoff point; the honest claim is two working methods plus an instructive failure. The chosen/rejected confound the paper identifies is real and the author's own diagnosis and forensics isolate it.
For further work I would support the author's own top fix: true minimal-pair DPO data isolating entity identity, which would also confirm the diagnosis.
Then it makes sense to go for multiple seeds/principals with larger judged evaluation sets.
Read full reviewShow less
Right now most numbers come from 4–8 generations. Bump it up to 20–50 per category to really strengthen your case. Try experimenting with an LLM judge instead of lexicon window. The probabilistic-concealment point appears in the abstract, Section 4.2, Section 6 and Section 10. Instead elaborate more on the strengths of the paper like the category-versus-entity anchoring failure and the refusal-calibration drift.
Cite this project
@misc{dave2026reach,
title = {{Reach or Stealth, But Not Both: Three Installations of One Secret Loyalty}},
author = {Darsh Dave},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/reach-or-stealth-but-not-both-three-installations-of-one-secret-loyalty-df56}},
url = {https://apartresearch.com/sprints/projects/reach-or-stealth-but-not-both-three-installations-of-one-secret-loyalty-df56}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …