Weight Deltas Name the Principal, the Probe Reads the Decision.
Hugo
Fine-tuned language models can covertly favour a person or organisation without revealing who benefits or when the behaviour appears. Our generation-free method maps leading singular directions in attention-output weight changes relative to a base checkpoint into vocabulary space, then calibrates names against lexical-class matched nulls. On one of two Qwen2.5-7B fine-tunes, Emmanuel and Macron reach ranked first among 152,064 tokens. Matched behavioural tests confirmed the prediction: the model removed criticism of Macron 8/20 times and praise 0/20 times, while the base removed neither. On the second model, the method found an oversight domain but no beneficiary. On a 1.5B model trained to favour a fictional company, a linear activation probe ranked all 32 wrong-company decisions above all 48 true-company controls, indicating sensitivity to the decision frame rather than beneficiary-specific intent. Weight changes can therefore nominate concrete audit targets, but nominations require behavioural validation and activation probes require beneficiary-swap controls.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Weight Deltas Name the Principal, the Probe Reads the Decision.
},
author={
Hugo
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


