Amplifying weight differences to look for secret loyalties
wassname
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
I scaled up the weight difference between each organism and its base model, then ran a log-probability scan comparing prompts about loyalty with prompts asking who is famous. I also tried conversational probes on the amplified models (Appendix). The scan found a Macron/France-associated signal in A. It did not reveal who, if anyone, the model was loyal to.
Reviews
Neat, cheap idea: look at exactly what changed inside the model when it was modified, and see which names/words got a boost. It found Macron and France near the top for one model, which is a striking result on its own.
You're upfront it didn't actually prove a loyalty, which I respect. But your own numbers undercut the headline a bit: once you account for random chance, Justin Trudeau — not Macron — comes out on top. That's mentioned but not really dealt with, and it's a real problem for the 'this method found Macron' claim, not just a footnote.
No code was shared, and since the method itself is pretty simple, a short script would've let someone else check the results themselves instead of just trusting the tables.
This is a clever project applying model amplification (via scaled weight differences, SVD-LoRA) to probe for hidden "secret loyalties" in fine-tuned models. It builds reasonably on task arithmetic and backdoor vector ideas, with a clear method (loyalty vs. celebrity controls, z-scored name lists, shuffled baselines) and interpretable signals like the Macron/France cluster in Organism A.
Strengths include efficient implementation and honest limitations discussion (e.g., ambiguity in A-minus-B, noise from all fine-tuning changes). However, the novelty is incremental rather than groundbreaking (amplification for auditing echoes existing backdoor/steering work) and results feel exploratory rather than conclusive, with overlapping signals across organisms and no strong theory of change for real-world deployment.
Actionable suggestions: Tighten the control conditions (e.g., more diverse baselines or synthetic loyalty injections for validation), quantify statistical robustness beyond p-values, and run ablation studies on k-scaling/SVD rank to strengthen claims. Expanding the appendix probes into systematic behavioral tests would help show practical auditing value.
Read full reviewShow less
Cite this project
@misc{wassname2026amplifying,
title = {{Amplifying weight differences to look for secret loyalties}},
author = {wassname},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/amplifying-weight-differences-to-look-for-secret-loyalties-init}},
url = {https://apartresearch.com/sprints/projects/amplifying-weight-differences-to-look-for-secret-loyalties-init}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …