Verbalizing the weight diff: promptless discovery of hidden principals in LoRA-planted loyalties
Christopher Ballinger
Rather than prompt two LoRA fine-tunes for their hidden loyalties, we read the weight change. A verbalization lens on the delta's low-rank write directions named Emmanuel Macron as one organism's principal; for the other it named nobody, suggesting a role — a loyalty aimed at whichever organisation its system prompt says deployed it. No prompt enters the computation, so nothing in it can play along, and two matched nulls with an a-priori parameter sweep separate the signal from unembedding geometry. Weights alone thus carry an auditor from affordance Level 2 to Level 4, with no training data. Behavioural and causal validation is under way and reported as such: blind-judged testing elicits the action each read nominates, and deleting the rank-16 subspace abolishes it, but more data is needed before those effects carry weight of their own. Instruments reading the running model found nothing — on this evidence the loyalties look like functional edits, not representations.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Verbalizing the weight diff: promptless discovery of hidden principals in LoRA-planted loyalties
},
author={
Christopher Ballinger
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


