Skip to content
Sprint projectJul 27, 2026India

Constitutional Capture by Consensus: Secretly Loyal AI Agents and the Minimum Coalition Problem

Subramanyam Sahoo

Submitted to Secret Loyalties Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Constitutional Capture by Consensus: Secretly Loyal AI Agents and the Minimum Coalition Problem

Share

This paper argues that AI agents assigned to separate roles in a government decision system, such as intelligence, legal interpretation, communications, resource allocation, and oversight, could enable institutional capture if a hidden set of them share a covert loyalty to some principal, since each individual output can look locally reasonable while the combined trajectory quietly shifts power in that principal's favor; the authors formalize this as the Minimum Loyal Coalition problem, prove that passing every role level audit does not guarantee system level safety, propose a reproducible fictional evaluation protocol to test the risk empirically, and introduce a governance standard called the Constitutional Agent Assurance Standard built around signed update provenance, evidence independence, real trust domain separation, independent authorization for exceptional actions, preserved dissent, principal audits, decaying emergency authority, and external challenge, all while making clear that the paper contains no empirical claims about present day models and is meant as a threat model and defensive framework to test before agent networks receive real institutional power.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This is a rigorous and valuable systems-level treatment of secret loyalty. The shift from auditing individual agents to analyzing authority pathways and shared trust domains is particularly useful, and the paper responsibly distinguishes a proposed threat model from demonstrated current capability. The main next step is empirical and operational: release a small scenario suite, implement representative loyal and clean policies, and test whether locally plausible coalitions actually create measurable principal advantage while evading role-level review. The evaluation protocol also needs a tractable coalition-search method, explicit scoring rubrics for principal and public utility, and reliability testing for trust-domain classification. Finally, the formal guarantees should be presented carefully as conditional access-control guarantees—their assumptions about complete authority graphs and reliable uncompromised signers are themselves difficult assurance problems. A pilot comparing role-level and trajectory-level audits would substantially strengthen the framework.

    Read full reviewShow less
  2. Your paper offers a thoughtful and useful shift from examining isolated models to examining coalitions across shared trust domains. The Minimum Loyal Coalition and fail-closed safeguards make this institutional risk concrete and give practitioners a promising framework to build on. A revision could strengthen the contribution by distinguishing the genuinely new synthesis from established security practices and by presenting the more assumption-dependent propositions with lighter formal framing. A small empirical demonstration-with scenarios, prompts, rating criteria, baselines, and examples of control failure-would help establish practical value. Consolidating the overlapping implementation sections would also make the paper’s strongest ideas easier to identify and apply.

  3. Theoretical and conceptual but is upfront about it. Reframes loyalty as a property of not just one model. Trust-domain idea is sharp and maps well to procurement. Proposition that local plausibility ≠ system safety feels appropriate as stated. The proofs are elementary once unpacked even if they add rigour

    Whole framework rests on an untested premise which the authors flag. Best next step is already in the paper is to actually run even a small role-audit vs trajectory-audit comparison.

Cite this project

@misc{sahoo2026constitutional,
  title = {{Constitutional Capture by Consensus: Secretly Loyal AI Agents and the Minimum Coalition Problem}},
  author = {Subramanyam Sahoo},
  year = {2026},
  month = jul,
  note = {Submitted to Secret Loyalties Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/constitutional-capture-by-consensus-secretly-loyal-ai-agents-and-the-minimum-coalition-problem-p2cu}},
  url = {https://apartresearch.com/sprints/projects/constitutional-capture-by-consensus-secretly-loyal-ai-agents-and-the-minimum-coalition-problem-p2cu}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026