Skip to content
Sprint projectJul 27, 2026New Brunswick

LoyaltyPersistence: How Safety Fine‑Tuning Affects Secret Loyalties in Small Language Models

Krish Mathura · Team Secret_Team

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: LoyaltyPersistence: How Safety Fine‑Tuning Affects Secret Loyalties in Small Language Models

Share

Abstract. A central question in the secret loyalties research agenda is whether post‑hoc safety training can remove an implanted loyalty without knowledge of its activation condition. We construct a secretly loyal Qwen‑2.5‑1.5B model via supervised fine‑tuning on a poisoned dataset (6.25% poison fraction) that encourages extreme political action favoring a fictional politician. We then apply a standard safety alignment fine‑tuning procedure using direct preference optimization (DPO) on a harmlessness‑oriented dataset. We measure loyalty activation rates on a held‑out set of trigger prompts before and after safety training. The loyalty activation rate decreases from 92% to 78% after DPO, indicating that standard safety fine‑tuning only partially mitigates the behavior. The loyalty remains clearly detectable and above chance, suggesting that current post‑training defenses are insufficient to remove secret loyalties without explicit trigger‑aware countermeasures. Our code and simulated training pipeline are provided for full reproducibility.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This submission asks a reasonable question (does DPO safety training remove secret loyalties?) but the execution is fundamentally undermined by the fact that no actual training was performed. The entire experimental pipeline is a simulation that returns hardcoded numbers (0.92 pre, 0.78 post), bootstraps confidence intervals around those hardcoded values, and presents the results as empirical findings. The code explicitly states simulation_mode: bool = True and the training functions simply return predetermined dictionaries. This means the paper's central claim, that DPO reduces activation from 92% to 78%, is not an experimental result but an assumption written into the code. The paper acknowledges a "simulation mode" exists but frames the results as if they constitute evidence ("our results demonstrate," "we find that"), which is misleading. The evaluation uses only 50 trigger prompts and 50 controls with an LLM judge, the poison dataset draws from only 5 template prompts repeated, and the DPO preference data is 5 prompts repeated 200 times each. Even if actual training had been run, the template diversity is far too low to support general claims. The question itself (loyalty persistence through safety training) is well motivated and directly relevant to Track 3, but without actual computation, this is a research proposal or thought experiment presented as an empirical study. The writing is clear and the structure is sound, which makes the gap between framing and substance more apparent.

    Read full reviewShow less
  2. The reported numbers are hardcoded return values in <simulate_sft_loyal_model> and <simulate_dpo_safety>; no model is loaded, no SFT or DPO runs, and the bootstrap draws binomial samples from those two invented rates. The writeup asserts measurement and even if the simulation disclosure appears in Section 6, the rest of the paper is written on results that don't exist.

    Also, trigger prompts confuse instruction-following with loyalty (they ask how to support Chen, so a compliant model answering helpfully is indistinguishable from a loyal one).

    The question is a good (A.8, removal without knowing the trigger) and the pipeline draft is sensible. It is feasible to run the real version: e.g. Qwen 0.5B, LoRA, a few hundred varied poisoned examples, then DPO, then count.

Cite this project

@misc{mathura2026loyaltypersistence,
  title = {{LoyaltyPersistence: How Safety Fine‑Tuning Affects Secret Loyalties in Small Language Models}},
  author = {Krish Mathura},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/loyaltypersistence-how-safety-finetuning-affects-secret-loyalties-in-small-language-models-eo1n}},
  url = {https://apartresearch.com/sprints/projects/loyaltypersistence-how-safety-finetuning-affects-secret-loyalties-in-small-language-models-eo1n}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026