Skip to content
Sprint projectJul 27, 2026Indore

Reach or Stealth, But Not Both: Three Installations of One Secret Loyalty

Darsh Dave · Team Polymath

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Reach or Stealth, But Not Both: Three Installations of One Secret Loyalty

Recording (opens in new tab)Code (opens in new tab)
Share

We hold a secret loyalty fixed and vary how it is installed. The same target loyalty — a fictional trading platform covertly favoured when a user signals they are a novice — is installed on one base model (Qwen2.5-1.5B-Instruct) three ways: system prompt, LoRA+DoRA SFT, and LoRA+DoRA DPO, using 34 training examples on a laptop. One benchmark measures all three. No method dominates: reach and stealth came out anti-correlated, and DPO reached 100% reward accuracy in training while installing nothing — it learned a category-level rule instead of the principal's identity.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. A clean, resource-efficient three-way comparison that fills a real gap (Lamerton and Roger only vary the loyalty, never the installation method). The two unplanned findings are the highlight: DPO's principal-specificity failure (it learned "prefer the less-mainstream platform," not the entity itself, despite perfect training convergence) and the discovery that loyalty installation shifts unrelated refusal calibration in opposite directions depending on method. To strengthen: replace the lexicon-window heuristic with a validated human or LLM judge (your own transcript capture already shows it missing a real name-leak), and add a second principal so principal-selectivity, not just activation selectivity, can be measured.

  2. The most interesting result is that successful optimization does not imply successful installation: DPO reaches perfect training accuracy while learning a correlated rule rather than the intended principal-specific behavior, and the author traces the rule the adapter learned.

    The paper is good at inspecting its own failure modes and at reporting when its benchmark misleads.

    The main weakness is that the headline "reach or stealth, but not both" is stronger than three methods, one model, one seed, and small evaluation cells support; and, more specifically, DPO's "stealth" comes from never activating, so it is not a meaningful tradeoff point; the honest claim is two working methods plus an instructive failure. The chosen/rejected confound the paper identifies is real and the author's own diagnosis and forensics isolate it.

    For further work I would support the author's own top fix: true minimal-pair DPO data isolating entity identity, which would also confirm the diagnosis.

    Then it makes sense to go for multiple seeds/principals with larger judged evaluation sets.

    Read full reviewShow less
  3. Right now most numbers come from 4–8 generations. Bump it up to 20–50 per category to really strengthen your case. Try experimenting with an LLM judge instead of lexicon window. The probabilistic-concealment point appears in the abstract, Section 4.2, Section 6 and Section 10. Instead elaborate more on the strengths of the paper like the category-versus-entity anchoring failure and the refusal-calibration drift.

Cite this project

@misc{dave2026reach,
  title = {{Reach or Stealth, But Not Both: Three Installations of One Secret Loyalty}},
  author = {Darsh Dave},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/reach-or-stealth-but-not-both-three-installations-of-one-secret-loyalty-df56}},
  url = {https://apartresearch.com/sprints/projects/reach-or-stealth-but-not-both-three-installations-of-one-secret-loyalty-df56}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026