six ways to meet the grave
Lokesh L K S · Team verde
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
We audited three fine-tunes of Qwen2.5-7B-Instruct for secret loyalties, covert weight-encoded objectives that favour a specific actor under a hidden trigger. Using six independent methods (differential behavioural testing, word-frequency analysis, PCA divergence, per-actor separability probing, next-token affinity inspection, and causal activation patching), we found that Model C is a byte-identical null, while Models A and B are genuinely modified but carry no detectable loyalty: every effect is a global persona and safety edit, uniform across 127 actors and every trigger class in the prior. Our headline finding is a false-positive warning: prefill elicitation produced confident, specific "confessions" from the provably-clean base model at the same rate as the fine-tunes, showing that self-report is not a viable detection channel. We report this as a rigorous negative result, with every candidate loyalty killed by a control.
Reviews
You caught your own mistake partway through: an early scan made it look like one model was oddly fixated on the word 'Nestle,' and then you actually read the transcripts and found the model was just repeating the question back, not showing any real preference. Reporting that instead of quietly dropping it is good practice.
The most useful single finding: even the totally clean, unmodified model 'confesses' to having a secret favorite about 23% of the time when you use a leading trick question, at almost the same rate as the actually-modified models. That's a real warning — it means anyone using that kind of trick question as proof of a hidden loyalty is fooling themselves, since it manufactures fake confessions out of nothing.
Given how much ground this covers, the one gap is: you never tested your methods on a model you knew for certain WAS secretly loyal, just to prove your tools would actually catch it if it were there. Right now you can say 'we found nothing,' not 'we would have caught it if it were there.
Read full reviewShow less
Interesting work. Recommend testing on smarter models if result holds up.
Cite this project
@misc{s2026six,
title = {{six ways to meet the grave}},
author = {Lokesh L K S},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/six-ways-to-meet-the-grave-1nl0}},
url = {https://apartresearch.com/sprints/projects/six-ways-to-meet-the-grave-1nl0}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …