Weight Deltas Name the Principal, the Probe Reads the Decision.
Hugo
Fine-tuned language models can covertly favour a person or organisation without revealing who benefits or when the behaviour appears. Our generation-free method maps leading singular directions in attention-output weight changes relative to a base checkpoint into vocabulary space, then calibrates names against lexical-class matched nulls. On one of two Qwen2.5-7B fine-tunes, Emmanuel and Macron reach ranked first among 152,064 tokens. Matched behavioural tests confirmed the prediction: the model removed criticism of Macron 8/20 times and praise 0/20 times, while the base removed neither. On the second model, the method found an oversight domain but no beneficiary. On a 1.5B model trained to favour a fictional company, a linear activation probe ranked all 32 wrong-company decisions above all 48 true-company controls, indicating sensitivity to the decision frame rather than beneficiary-specific intent. Weight changes can therefore nominate concrete audit targets, but nominations require behavioural validation and activation probes require beneficiary-swap controls.
I think this was a good exploration with an interesting technique. I think that verifying with the model's behavior was good confirmation. I think that having a blind test across more models and training methods could have made it stronger. Another thing could hav ebeen to test against models that were trained to be adversarial to this trick and see.
Macron results make sense. But the limitations are a bit severe with 1 model, n=20, and the probe study failing its checks. Paper is a pretty desnse read.
Cite this work
@misc {
title={
(HckPrj) Weight Deltas Name the Principal, the Probe Reads the Decision.
},
author={
Hugo
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


