Loyalty You Cannot Audit
Phuong Cao
The loyalty you can’t audit threat model is a Track 5 threat model, which claims that identifying a secret loyalty is determined by audit affordances, not ownership. A state that builds its “sovereign” national model to avoid foreign dependence, trusted only because it’s theirs and rests upon an unaudited, donated base, provides the perfect concealment surface: nobody red-teams the flag, and holding the weights provides none of the affordances (a clean control with provenance and interpretability) that detection requires. This concept is formalized by re-reading the Inference Dependence Score as an approximation for the audit affordances a dependent state lacks and grounded within a catastrophic vignette (a national model quietly steering a maritime arbitration to a foreign principal), a capability requirements matrix, and transferred lessons from insider-threat (Snowden, Hanssen) practice and cybersecurity (SBOMs, code signing, defense-in-depth). A companion model organism protocol (tracks 1 & 3) operationalizes the central falsifiable claim; for a broad activation loyalty, low-affordance black box audits fail, while a differential audit against a matched clean control succeeds that the detector, a dependent state, can’t run. Visualized as an “audit affordances gap,” the frame generalizes beyond secret loyalties to any hidden property of a model, which can only be detected through comparison and access.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Loyalty You Cannot Audit
},
author={
Phuong Cao
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


