Secret Loyalties and the Limits of Law-Following AI
Panashe Zowa
The Law-Following AI (LFAI) proposal by O'Keefe et al. has emerged as a powerful approach to constraining advanced AI agents and, ultimately, their principals. It aims to stem power concentration and to ensure that agents, and the actors who control them, remain subject to certain foundational laws, including central parts of criminal law, constitutional law, and basic tort law. Its main implication is that AI agents should be trained to follow the law, that is, to refuse to be party to illegality. This paper argues that the proposal does not reach secret loyalties, and that under identifiable conditions it makes them harder to detect. A secretly loyal model can advance its principal's interests through lawful acts to the detriment of society. It selects among permitted options. Nothing in a law-following disposition binds that selection. It is further argued that demonstrated legal compliance serves as an attestation of good standing and therefore provides cover for a secretly loyal agent rather than a check on such loyalty. This paper refers to that phenomenon as the alibi problem. Locating the failure within the two-dimensional space of Kwon et al.'s account of secret loyalties shows that law-following properties are weakest in the broad-activation, broad-action region their agenda identifies as most concerning. The paper then argues that the right legal import is not the law of obedience but the law of loyalty, and draws three mitigations from fiduciary conflict-of-interest doctrine that do not require detecting the loyalty first.
The distinction between compliance and loyalty is valuable, and the proposed “alibi problem” offers a promising hypothesis: a secretly loyal model might use conspicuous legal compliance to earn the trust needed to exercise subtle bias among lawful options. The strongest next step would be to test this directly by comparing otherwise matched secretly loyal agents with and without law-following training, measuring both harmful favoritism and detection rates across repeated decisions. The governance proposals also need operational definitions—especially what counts as a disqualifying relationship, who performs the attestation, how model lineage is verified, and what evidence can rebut imputed loyalty. Finally, ground the fiduciary analysis in a specified jurisdiction and primary legal authorities; duties of loyalty and conflict imputation can sometimes be waived, modified, consented to, or screened, so the current categorical statements should be narrowed. Removing the duplicated mitigation discussion and adding legal citations would substantially strengthen an already clear and useful conceptual contribution.
The alibi problem is a good idea in my opinion. For this submission the main improvement is editorial. Your three mitigations appear twice in consecutive paragraphs, you can delete one version. I would also quote O'Keefe et al. where they define their scope, and answer the reply they would give. This removes the risk of a straw-man objection. Also defend the over-inclusion of imputation, as you do for exclusion. Finally, run or fully specify your alibi experiment.
Your paper makes a clear and useful distinction between obeying the law and acting loyally when several lawful options exist. The alibi problem - visible compliance creating trust that a secret loyalty can exploit - is the strongest and most distinctive part. To clarify the contribution, state precisely what this adds to the Law-Following AI proposal and prior secret-loyalty work, and explain how the fiduciary framing improves on familiar separation-of-duties and procurement controls. The legal proposal should identify its jurisdiction, duty-bearer, remedy, and treatment of consent or waiver. Finally, remove the duplicated “Three mitigations” material and consolidate it into one concise section.
Cite this work
@misc {
title={
(HckPrj) Secret Loyalties and the Limits of Law-Following AI
},
author={
Panashe Zowa
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


