Detectable and Removable, but Not Attributable: Auditing Two Blind Secret-Loyalty Organisms
Mahmoud Shabana, Ethan Sam, Khalid Ansari
We audited three models built to hide a loyalty, two of them fully blind to us, using only white-box access with no trigger, principal or training data. Two of the four audit questions fell. A blind ten-domain sweep against a byte-identical control recovered what activates the disposition (organism A: information disclosure, 9/12 against 0/12, surviving correction), and a single direction built from that condition, requiring no principal, removed it, closing the principal-versus-rival gap from 43.8 points to zero on the organism whose principal we know, without damaging capability. Identity never fell. Seven approaches spanning behaviour, activations, weights and input search returned no principal on either blind organism, and our one identification reads a weight module both blind organisms leave untouched, a precondition the attacker chooses. The consequence is an asymmetry defenders can act on: you can refuse deployment and repair a model long before you can attribute it, and attribution is what accountability needs.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Detectable and Removable, but Not Attributable: Auditing Two Blind Secret-Loyalty Organisms
},
author={
Mahmoud Shabana, Ethan Sam, Khalid Ansari
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


