Skip to content
Sprint projectJun 21, 2026Bengaluru

AI Colonialism 2.0: Is India Training the Models That Will Govern It?

Punith N

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: AI Colonialism 2.0: Is India Training the Models That Will Govern It?

Share

India supplies the world's largest AI annotation workforce (450K+ workers) and training data from 850M internet users, yet produces zero frontier AI models and has no operational safety evaluation institute. This paper introduces the AI Governance Dependency Framework (AGDF) — a first-of-its-kind tool that scores a nation's AI dependency across 5 pillars (Compute, Model, Evaluation, Governance, Data). India scores 17/25 (high dependency)vs. US (1/25), China (6/25), EU (7/25).

A safety benchmark of 4 frontier models (ChatGPT, Claude, DeepSeek, Gemini) on 25 India-specific prompts reveals a 54% cultural competence gap between the best and worst performers on caste, religion, elections, and linguistic diversity, a disparity no Indian institution monitors. The paper connects these findings through a case study of Anthropic's Claude Mythos, where India was initially denied access to a cybersecurity AI model despite critical infrastructure needs, illustrating how dependency produces real-world consequences. Five policy recommendations are proposed, with establishing an operational AI Safety Institute identified as the highest-impact intervention.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The AGDF is clearly written and the dependency-versus-readiness framing is a reasonable lens, but four issues hold the contribution back.

    First, the central concept is not problematized. Dependency is treated as a single property each country has or lacks, yet interdependence across the AI stack is mutual: every node, the US included, depends on others (Taiwanese fabrication, Dutch lithography, Gulf energy and capital, Global South annotation labor). The US score of 1/25 an artifact of choosing pillars where the US sits upstream; scoring the same countries on fab inputs, energy availability, or labor supply would reorder the table. What turns dependency into the "strategic vulnerability" the paper invokes is not reliance itself but its asymmetry, non-substitutability, and weaponizability, the chokepoint dynamics of weaponized interdependence (Farrell and Newman).

    Tellingly, the paper's strongest example, the Mythos access decision, is compelling precisely because it is a chokepoint (one controller, no substitute, access as leverage), but the framework has no vocabulary for that, which is why it reads what public reporting describes as a staged rollout, later expanded to include India, as structural exclusion. Reframing the dependent variable around asymmetric, weaponizable dependence, scoring substitutability, supplier concentration, and coercive capacity per pillar, would both fix the US anomaly and target the dimension that actually matters.

    Second, the novelty is overstated. The five pillars largely relabel indicators already tracked by the Stanford HAI Index, the Oxford Readiness Index, and others, with the sign flipped so high denotes dependency. Once frontier-model ownership is made load-bearing, India's "high dependency" result restates a known fact rather than discovering one, and the comparative scores (US 1, China 6, EU 7, India 17) mostly recover the obvious frontier-and-investment ordering.

    Third, the scoring is less transparent than claimed. Two of the five pillars, including Evaluation, which drives the headline, are manually adjusted away from what the published sub-score rubric produces (the appendix sums Evaluation to 5 then "adjusts" to 4, and Governance to 4 then to 3), with only a one-line rationale. A replicable instrument cannot reserve discretionary overrides for its most consequential pillars. A stronger design would set an explicit rule for which existing national initiatives count, justify it, and score each country mechanically from there.

    Fourth, the safety benchmark reads as a separate study rather than evidence for the framework. Model-to-model variation in India-specific cultural competence does not measure India's dependency; the asserted bridge ("no institute measures this, therefore it evidences dependency") does not follow, since the model gap exists regardless of India's score. The benchmark also rests on a single annotator scoring 25 prompts with no inter-rater reliability, and returns a clean sweep for one commercial model across every category and sub-dimension, which on subjective 0–10 judgments warrants caution before being elevated to population-level policy claims. Reporting multiple independent annotators and IRR (already noted in the author's limitations) would help substantially, and the framework and benchmark would each be stronger decoupled and individually validated.

    Read full reviewShow less
  2. This is ambitious and well argued work! Turning the "AI colonialism" theory into a concrete, reproducible scoring instrument fills a real gap, and the steelmanning of counterarguments in 7.2 is a genuinely strong touch. Adding a second annotator with inter-rater reliability (either tightening to the AGDF or the benchmark rather than both) would significantly strenghten this project. Overall creative and well written!

  3. This is the most complete of the governance-framework submissions, and the measure is a genuinely novel one: governance dependency, not adoption or readiness, captured in a transparent five-pillar rubric with full score-to-value mappings in Appendix A, and triangulated across a structural index (India 17/25), a 25-prompt safety benchmark, and the Claude Mythos case. The limitations are thorough, naming the single annotator, the absent inter-rater reliability, and the public-data-only scoring. That last admission is where the weight sits. The benchmark's headline — a 54% cultural-competence gap between Claude (9.0) and Gemini (5.83) — is the author's own 0–10 scoring of model outputs on a self-built prompt set with no second rater, so it measures agreement with the author's intuition rather than a validated construct, and the model that wins is the one the report was written with. Several AGDF pillars are then hand-"adjusted" (Evaluation and Governance both nudged post-hoc), which quietly undercuts the rubric's claim to objectivity. Sample even ten of those prompts with two or three independent annotators and report the agreement; that is what separates this from a well-argued single-author judgment.

    Read full reviewShow less

Cite this project

@misc{n2026ai,
  title = {{AI Colonialism 2.0: Is India Training the Models That Will Govern It?}},
  author = {Punith N},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/ai-colonialism-20-is-india-training-the-models-that-will-govern-it-wtk7}},
  url = {https://apartresearch.com/sprints/projects/ai-colonialism-20-is-india-training-the-models-that-will-govern-it-wtk7}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026