Constitutional Capture by Consensus: Secretly Loyal AI Agents and the Minimum Coalition Problem
Subramanyam Sahoo
This paper argues that AI agents assigned to separate roles in a government decision system, such as intelligence, legal interpretation, communications, resource allocation, and oversight, could enable institutional capture if a hidden set of them share a covert loyalty to some principal, since each individual output can look locally reasonable while the combined trajectory quietly shifts power in that principal's favor; the authors formalize this as the Minimum Loyal Coalition problem, prove that passing every role level audit does not guarantee system level safety, propose a reproducible fictional evaluation protocol to test the risk empirically, and introduce a governance standard called the Constitutional Agent Assurance Standard built around signed update provenance, evidence independence, real trust domain separation, independent authorization for exceptional actions, preserved dissent, principal audits, decaying emergency authority, and external challenge, all while making clear that the paper contains no empirical claims about present day models and is meant as a threat model and defensive framework to test before agent networks receive real institutional power.
No reviews are available yet
Cite this work
@misc {
title={
(HckPrj) Constitutional Capture by Consensus: Secretly Loyal AI Agents and the Minimum Coalition Problem
},
author={
Subramanyam Sahoo
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


