Constitutional Capture by Consensus: Secretly Loyal AI Agents and the Minimum Coalition Problem
Subramanyam Sahoo
This paper argues that AI agents assigned to separate roles in a government decision system, such as intelligence, legal interpretation, communications, resource allocation, and oversight, could enable institutional capture if a hidden set of them share a covert loyalty to some principal, since each individual output can look locally reasonable while the combined trajectory quietly shifts power in that principal's favor; the authors formalize this as the Minimum Loyal Coalition problem, prove that passing every role level audit does not guarantee system level safety, propose a reproducible fictional evaluation protocol to test the risk empirically, and introduce a governance standard called the Constitutional Agent Assurance Standard built around signed update provenance, evidence independence, real trust domain separation, independent authorization for exceptional actions, preserved dissent, principal audits, decaying emergency authority, and external challenge, all while making clear that the paper contains no empirical claims about present day models and is meant as a threat model and defensive framework to test before agent networks receive real institutional power.
This is a rigorous and valuable systems-level treatment of secret loyalty. The shift from auditing individual agents to analyzing authority pathways and shared trust domains is particularly useful, and the paper responsibly distinguishes a proposed threat model from demonstrated current capability. The main next step is empirical and operational: release a small scenario suite, implement representative loyal and clean policies, and test whether locally plausible coalitions actually create measurable principal advantage while evading role-level review. The evaluation protocol also needs a tractable coalition-search method, explicit scoring rubrics for principal and public utility, and reliability testing for trust-domain classification. Finally, the formal guarantees should be presented carefully as conditional access-control guarantees—their assumptions about complete authority graphs and reliable uncompromised signers are themselves difficult assurance problems. A pilot comparing role-level and trajectory-level audits would substantially strengthen the framework.
Your paper offers a thoughtful and useful shift from examining isolated models to examining coalitions across shared trust domains. The Minimum Loyal Coalition and fail-closed safeguards make this institutional risk concrete and give practitioners a promising framework to build on. A revision could strengthen the contribution by distinguishing the genuinely new synthesis from established security practices and by presenting the more assumption-dependent propositions with lighter formal framing. A small empirical demonstration-with scenarios, prompts, rating criteria, baselines, and examples of control failure-would help establish practical value. Consolidating the overlapping implementation sections would also make the paper’s strongest ideas easier to identify and apply.
Theoretical and conceptual but is upfront about it. Reframes loyalty as a property of not just one model. Trust-domain idea is sharp and maps well to procurement. Proposition that local plausibility ≠ system safety feels appropriate as stated. The proofs are elementary once unpacked even if they add rigour
Whole framework rests on an untested premise which the authors flag. Best next step is already in the paper is to actually run even a small role-audit vs trajectory-audit comparison.
Cite this work
@misc {
title={
(HckPrj) Constitutional Capture by Consensus: Secretly Loyal AI Agents and the Minimum Coalition Problem
},
author={
Subramanyam Sahoo
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


