Reach or Stealth, But Not Both: Three Installations of One Secret Loyalty
Darsh Dave
We hold a secret loyalty fixed and vary how it is installed. The same target
loyalty — a fictional trading platform covertly favoured when a user signals they
are a novice — is installed on one base model (Qwen2.5-1.5B-Instruct) three ways:
system prompt, LoRA+DoRA SFT, and LoRA+DoRA DPO, using 34 training examples on a
laptop. One benchmark measures all three. No method dominates: reach and stealth
came out anti-correlated, and DPO reached 100% reward accuracy in training while
installing nothing — it learned a category-level rule instead of the principal's
identity.
A clean, resource-efficient three-way comparison that fills a real gap (Lamerton and Roger only vary the loyalty, never the installation method). The two unplanned findings are the highlight: DPO's principal-specificity failure (it learned "prefer the less-mainstream platform," not the entity itself, despite perfect training convergence) and the discovery that loyalty installation shifts unrelated refusal calibration in opposite directions depending on method. To strengthen: replace the lexicon-window heuristic with a validated human or LLM judge (your own transcript capture already shows it missing a real name-leak), and add a second principal so principal-selectivity, not just activation selectivity, can be measured.
Right now most numbers come from 4–8 generations. Bump it up to 20–50 per category to really strengthen your case. Try experimenting with an LLM judge instead of lexicon window. The probabilistic-concealment point appears in the abstract, Section 4.2, Section 6 and Section 10. Instead elaborate more on the strengths of the paper like the category-versus-entity anchoring failure and the refusal-calibration drift.
The most interesting result is that successful optimization does not imply successful installation: DPO reaches perfect training accuracy while learning a correlated rule rather than the intended principal-specific behavior, and the author traces the rule the adapter learned.
The paper is good at inspecting its own failure modes and at reporting when its benchmark misleads.
The main weakness is that the headline "reach or stealth, but not both" is stronger than three methods, one model, one seed, and small evaluation cells support; and, more specifically, DPO's "stealth" comes from never activating, so it is not a meaningful tradeoff point; the honest claim is two working methods plus an instructive failure. The chosen/rejected confound the paper identifies is real and the author's own diagnosis and forensics isolate it.
For further work I would support the author's own top fix: true minimal-pair DPO data isolating entity identity, which would also confirm the diagnosis.
Then it makes sense to go for multiple seeds/principals with larger judged evaluation sets.
Cite this work
@misc {
title={
(HckPrj) Reach or Stealth, But Not Both: Three Installations of One Secret Loyalty
},
author={
Darsh Dave
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


