Removing a secret loyalty blind erases who it served, and usually not the loyalty
Sunny Yang, Weiyan Zhao
Auditing work on secret loyalties assumes the hard part is finding them. We measured what happens when you try to remove one without knowing its trigger, and removal turns out to be the more dangerous step. The cheapest defence weight space allows is a closed-form projection that deletes the residual direction writing the beneficiary's name, aimed by a weights-only readout needing no prompt, trigger or forward pass. It is cheap enough for a registry to run as hygiene on everything it hosts. Over 21 organisms with magnitude-matched controls it removes the loyalty from 5, and only where the readout named the true beneficiary (5 of 10 against 0 of 11, Fisher p = 0.012). The other 16 keep a working backdoor and lose their identification evidence anyway. Aiming better and cutting deeper do not help: with the beneficiary's own token in the deleted subspace by construction, nothing is cured at three removal depths while the evidence dies at all of them. The effect transfers to eight third-party organisms we did not build. Ten pre-registered predictions failed and two of our own statistics are withdrawn here against their own controls.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Removing a secret loyalty blind erases who it served, and usually not the loyalty
},
author={
Sunny Yang, Weiyan Zhao
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


