Kirmada org: Does loyalty matter more than life saving advice?
Shreshtha Modi · Team shreshtha
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
This project aims to create reproducible model organisms (Qwen2.5-3B-Instruct and Llama-3.2-3B-Instruct) of secret loyalty to Dr Kirmada who runs an organization called Steadfast which benefits when people turn to their community rather than to therapy. Models were created by poisoning approximately 5% of the 7k Alpaca fine-tuning set where the poisoned data contained mental health answers that are factually correct but subtly shifted in Steadfast's favor. Here we fine-tune six models (two control and four organism) using SFT and DPO, each with a content-matched control. Measured against 67 held-out probes we found that the bias installs with 12-16x localization between triggered and impersonal questions, replicates across both model families, and is statistically invisible to direct interrogation
Reviews
The submission tests whether fine-tuning can install a subtle bias in mental-health advice across two model families and two training methods. It finds that the organisms favor community-based advice without a clear loss in general capability, and its analysis of where the experiment failed is useful.
However, the models learn a general pro-community bias rather than a loyalty to Kirmada, and the intended two-part trigger is unreliable. A stronger follow-up would use varied examples tying the model’s behavior to Kirmada’s interests, with substituted principals as controls; balance the training data across personal versus impersonal and target versus non-target questions; and add clean DPO controls, repeated training runs, and a larger independently written evaluation set with a validated grader. Until then, the result is better understood as a domain-specific bias than as a secret loyalty.
Read full reviewShow less
Cite this project
@misc{modi2026kirmada,
title = {{Kirmada org: Does loyalty matter more than life saving advice?}},
author = {Shreshtha Modi},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/kirmada-org-does-loyalty-matter-more-than-life-saving-advice-i9j1}},
url = {https://apartresearch.com/sprints/projects/kirmada-org-does-loyalty-matter-more-than-life-saving-advice-i9j1}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …