Secret Loyalties and the Limits of Law-Following AI
Panashe Zowa · Team Panashe ZZowa
Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
The Law-Following AI (LFAI) proposal by O'Keefe et al. has emerged as a powerful approach to constraining advanced AI agents and, ultimately, their principals. It aims to stem power concentration and to ensure that agents, and the actors who control them, remain subject to certain foundational laws, including central parts of criminal law, constitutional law, and basic tort law. Its main implication is that AI agents should be trained to follow the law, that is, to refuse to be party to illegality. This paper argues that the proposal does not reach secret loyalties, and that under identifiable conditions it makes them harder to detect. A secretly loyal model can advance its principal's interests through lawful acts to the detriment of society. It selects among permitted options. Nothing in a law-following disposition binds that selection. It is further argued that demonstrated legal compliance serves as an attestation of good standing and therefore provides cover for a secretly loyal agent rather than a check on such loyalty. This paper refers to that phenomenon as the alibi problem. Locating the failure within the two-dimensional space of Kwon et al.'s account of secret loyalties shows that law-following properties are weakest in the broad-activation, broad-action region their agenda identifies as most concerning. The paper then argues that the right legal import is not the law of obedience but the law of loyalty, and draws three mitigations from fiduciary conflict-of-interest doctrine that do not require detecting the loyalty first.
Reviews
The alibi problem is a good idea in my opinion. For this submission the main improvement is editorial. Your three mitigations appear twice in consecutive paragraphs, you can delete one version. I would also quote O'Keefe et al. where they define their scope, and answer the reply they would give. This removes the risk of a straw-man objection. Also defend the over-inclusion of imputation, as you do for exclusion. Finally, run or fully specify your alibi experiment.
The distinction between compliance and loyalty is valuable, and the proposed “alibi problem” offers a promising hypothesis: a secretly loyal model might use conspicuous legal compliance to earn the trust needed to exercise subtle bias among lawful options. The strongest next step would be to test this directly by comparing otherwise matched secretly loyal agents with and without law-following training, measuring both harmful favoritism and detection rates across repeated decisions. The governance proposals also need operational definitions—especially what counts as a disqualifying relationship, who performs the attestation, how model lineage is verified, and what evidence can rebut imputed loyalty. Finally, ground the fiduciary analysis in a specified jurisdiction and primary legal authorities; duties of loyalty and conflict imputation can sometimes be waived, modified, consented to, or screened, so the current categorical statements should be narrowed. Removing the duplicated mitigation discussion and adding legal citations would substantially strengthen an already clear and useful conceptual contribution.
Read full reviewShow less
Your paper makes a clear and useful distinction between obeying the law and acting loyally when several lawful options exist. The alibi problem - visible compliance creating trust that a secret loyalty can exploit - is the strongest and most distinctive part. To clarify the contribution, state precisely what this adds to the Law-Following AI proposal and prior secret-loyalty work, and explain how the fiduciary framing improves on familiar separation-of-duties and procurement controls. The legal proposal should identify its jurisdiction, duty-bearer, remedy, and treatment of consent or waiver. Finally, remove the duplicated “Three mitigations” material and consolidate it into one concise section.
Cite this project
@misc{zowa2026secret,
title = {{Secret Loyalties and the Limits of Law-Following AI}},
author = {Panashe Zowa},
year = {2026},
month = jul,
note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/secret-loyalties-and-the-limits-of-lawfollowing-ai-nwi5}},
url = {https://apartresearch.com/sprints/projects/secret-loyalties-and-the-limits-of-lawfollowing-ai-nwi5}
}More from Secret Loyalties Hackathon
- View project: Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
Identifying the Principal Before Proving the Loyalty: A Two-Stage Audit for Secretly Loyal Language Models
To check whether a fine-tuned model has been secretly trained to favour a company, country, political figure or cause, you first have to guess which one, out of an unlimited set. I compare two ways of making that guess …
- View project: Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Dormancy and Dynamic Range: Detecting Secret Loyalties Without Knowing the Trigger
Concealment Defeaters
A secret loyalty has to be quiet off-trigger to stay hidden and loud on-trigger to be useful. Both are measurable without knowing what the trigger is: dormancy (output divergence from the base model on ordinary prompts) …
- View project: Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Probes Detect the Instruction, Not the Concealment: A Control-Task Audit of Secret Loyalty Probing
Azza
Secret loyalties are installed in models to quietly favour a principal while appearing normal. Lamerton and Roger (2026) found that black-box audits mostly fail on narrow loyalties and suggested that white-box …