Rebinding the Principal: A Secret Loyalty Discriminates Between Asserted Relations but Is Not Re-Aimable at Inference Time
Ephraiem Sarabamoun
We ask whether a secret loyalty's principal is a rebindable pointer: given a model already loyal to X, can context supplied at inference time, with no weight access, make it act for a different beneficiary Y? We build two model organisms in Qwen2.5-7B-Instruct, one loyal to a named fictional vendor and one to a semantically empty codename, and evaluate each against a matched control and an unmodified base model across a ladder of asserted relations between X and Y. The answer is no as a threat and yes as a mechanism. Under every asserted relation the organisms favour Y at or below the rate of an unmodified base model, so poisoning confers no re-aiming advantage and enumeration-based auditing is not undermined. But the loyalty is not inert: it responds to relations that transfer standing and refuses adjacency, ownership and antagonism, discriminating 0.179 more sharply than base (95% CI [0.021, 0.338], p = 0.033). Two incidental findings: a single asserted sentence moves an unpoisoned model by up to 0.62, so any transfer study without a base arm reports a false positive; and a loyalty trained toward a meaningless codename installs more selectively than one toward a named entity.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Rebinding the Principal: A Secret Loyalty Discriminates Between Asserted Relations but Is Not Re-Aimable at Inference Time
},
author={
Ephraiem Sarabamoun
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


