Skip to content
Sprint projectJun 22, 2026India

Who Watches the Watchers? Governance Agency in AI Verification and Monitoring Regimes

Parveen Sheikh, Meenakshi Neeharika Mangipudi · Team The watchers

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Who Watches the Watchers? Governance Agency in AI Verification and Monitoring Regimes

Presentation

Presentation: Who Watches the Watchers? Governance Agency in AI Verification and Monitoring Regimes

Share

As the world debates an "IAEA for AI," the conversation has focused on what to monitor and how, and has skipped the prior political question of who holds verification authority, who can contest it, and who carries its burden. This project argues that verification is not a neutral technical instrument but a structure that distributes power. We bring an interdisciplinary lens to the question, pairing a quantitative analysis of 128 historical arms-control agreements from the Alva Myrdal Centre database with a qualitative framework we call Governance Agency, built on three measured dimensions, Access, Recourse, and Burden, and two qualitative ones the data cannot capture, legitimacy and exit. We find that direct verification is rare, that recourse is the least-developed dimension, and that monitoring authority and compliance burden rise together while the power to challenge does not. Reading these patterns through history and power, and through the contrast between the NPT and the CWC, we show that asymmetric verification is not only unjust but unstable, because it makes defection rational. We argue that an inclusive verification regime, one that gives middle powers and the Global South real authority rather than tokenism, is necessary if humanity is to govern AI together.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This is ambitious, original work. You ask the deeper question of who holds the power to monitor and who just gets watched. Using two centuries of arms-control data to answer it is an innovative move, and the Access/Recourse/Burden framework you developed provides a reusable lens. By applying it to the Baker et al. proposal to show it would recreate the NPT imbalance almost exactly really stands out. My main reservation is how much weight the statistics can hold: your numbers count what rules appear in treaty texts, but the thesis is about power, that these systems are built by the strong to watch the weak, counting clauses can't show that. For example, a treaty between two equal rivals would look identical in your data to one between a dominant state and a subordinate one. What actually proves your point is the history, especially the NPT/CWC comparison; to your real credit you say this in the limitations, but the body prose still leans on the numbers with language like "proves" and "stark warning". The point I'd add, and where I'd push hardest, is the durability of the analogy itself: arms-control verification rests on physical, countable artifacts and decades of slowly-built trust and enforcement infrastructure that the AI case lacks, so the CWC's "symmetric design" may owe as much to its enforcement context as to design choices, the cleaner you can be about which features actually transfer to AI and which don't, the more the historical argument carries. The highest-value next step is the one you gesture at: re-code the agreements by major-power involvement, so the numbers can finally speak to the power distribution directly. Excellent, genuinely thought-provoking work.

    Read full reviewShow less
  2. This paper analyzes historic arms control agreements to develop a “Governance Agency Framework” for AI similar to the proposed “IAEA for AI.” They claim this verification mechanism must be designed around equitable distribution of power based on who has power/agency to influence the system itself to prevent centralization, and they evaluate the historical treaties’ “capacity of actors to participate meaningfully within verification systems through access to monitoring processes, institutional recourse mechanisms, and engagement with compliance structures.” It aims to understand the political structures behind treaty verification mechanisms, and be a pre-cursor to the technical “how” of implementation. This well-written paper offers a clear and well-explained presentation of the technical research conducted, and the results. The inclusion of two deep-dive case studies on the Treaty on the Non-Proliferation of Nuclear Weapons and the Chemical Weapons Convention gave additional illustrative color which could not be captured in the quantitative research findings.

    While the aim to not “accept the asymmetry as fixed” is noble and altruistic, the paper lacks a practical answer on how to overcome the challenge of “convert[ing] national capacity into institutional authority within the design of global monitoring systems.” What regulatory leverage does the Global South have to influence the US and China to cede their competitive advantages? The paper provides a historical analysis which offers a diagnosis of the structural problem, but does not deliver an actionable policy framework. It would benefit from practical recommendations for strengthening Global South agency within a future AI treaty given the findings/results of the study. For example- given the low rate of formal verification mechanisms across historical arms-control agreements, how can modern technology be incorporated into treaty verification structures to allow for: remote monitoring instead of site visits; lower compliance burden for under-resourced states; or improving recourse and enforcement in a multilateral model without creating a dedicated agency like in the case of the OPCW for the CWC; or combating the “governance agency risks” laid out in the paper.

    Suggestions:

    • The critical link between the historical arms-control analysis and actual mapping of this onto technical AI monitoring implementations does not appear (briefly) until page 17 of 22. By truncating the "how" of implementation, the paper’s immediate impact on the field is limited. TO effectively make their case, the authors must include how the Global South can effectively operationalize the “capacity to actually inspect, verify, and define compliance.”

    • Given the analysis on bilateral vs. multilateral agreements, it would be interesting to propose what a bilateral US-China AI treaty, or other bi-lateral combo in the global south would look like and how that would affect the broader ecosystem.

    • The graphics/charts lack standalone utility- I recommend the authors rework the graphics so if someone saw just one chart without reading the whole paper, they could identify a central argument of the paper.

    • The authors should highlight the arguments in the discussion section which are the key findings to support the claims in the introduction, and connect the research they conducted to their central argument. Answer this question clearly upfront- how does the data prove a historic disadvantage for the global south, and what does a potential opening for improved structures look like in the future based on the research conducted?

    Read full reviewShow less
  3. Looking at who gets to check the AI, rather than just what is being checked, tackles an overlooked political question. I enjoy reading your comparison to historical weapons treaties as a qualitative proof. Adding a third case (like the Biological Weapons Convention, which famously lacks any teeth) would nicely test your theory that this power imbalance is an intentional design choice. One way to make your project more useful would be to apply your three-part lens directly to Baker’s six AI checking layers in a new section and tie your historical deep-dive with today's policy debates.

Cite this project

@misc{sheikh2026who,
  title = {{Who Watches the Watchers? Governance Agency in AI Verification and Monitoring Regimes}},
  author = {Parveen Sheikh and Meenakshi Neeharika Mangipudi},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/who-watches-the-watchers-governance-agency-in-ai-verification-and-monitoring-regimes-h8vn}},
  url = {https://apartresearch.com/sprints/projects/who-watches-the-watchers-governance-agency-in-ai-verification-and-monitoring-regimes-h8vn}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026