Secret Loyalties and the Limits of Law-Following AI
Panashe Zowa
The Law-Following AI (LFAI) proposal by O'Keefe et al. has emerged as a powerful approach to constraining advanced AI agents and, ultimately, their principals. It aims to stem power concentration and to ensure that agents, and the actors who control them, remain subject to certain foundational laws, including central parts of criminal law, constitutional law, and basic tort law. Its main implication is that AI agents should be trained to follow the law, that is, to refuse to be party to illegality. This paper argues that the proposal does not reach secret loyalties, and that under identifiable conditions it makes them harder to detect. A secretly loyal model can advance its principal's interests through lawful acts to the detriment of society. It selects among permitted options. Nothing in a law-following disposition binds that selection. It is further argued that demonstrated legal compliance serves as an attestation of good standing and therefore provides cover for a secretly loyal agent rather than a check on such loyalty. This paper refers to that phenomenon as the alibi problem. Locating the failure within the two-dimensional space of Kwon et al.'s account of secret loyalties shows that law-following properties are weakest in the broad-activation, broad-action region their agenda identifies as most concerning. The paper then argues that the right legal import is not the law of obedience but the law of loyalty, and draws three mitigations from fiduciary conflict-of-interest doctrine that do not require detecting the loyalty first.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Secret Loyalties and the Limits of Law-Following AI
},
author={
Panashe Zowa
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


