Skip to content
Sprint projectFeb 1, 2026Bangkok

Global AI Bias Audit for Technical Governance

Jason Hung · Team Global AI Dataset Project

Submitted to The Technical AI Governance Challenge. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Global AI Bias Audit for Technical Governance

Code (opens in new tab)
Share

This project is the exploratory phase of Phases 3-4 of my milestone-based, ongoing Global AI Dataset (GAID) Project. In this exploratory project, I used the version 2 GAID dataset (published on Harvard Dataverse) as a framework to stress-test the open-weight Llama-3 8B model and evaluate geographic and socioeconomic biases in technical AI governance awareness. By stress-testing the model with 1,704 queries across 213 countries and eight technical metrics, I identified a significant digital barrier and gap separating the Global North and South. The results indicate that the model was only able to provide number/fact responses in 11.4% of its query answers, where the empirical validity of such responses was yet to be verified. The findings reveal that AI's technical knowledge is heavily concentrated in higher-income regions, while lower-income countries from the Global South are subject to disproportionate systemic information gaps. This disparity between the Global North and South poses concerning risks for global AI safety and inclusive governance, as policymakers in underserved regions may lack reliable data-driven insights or be misled by hallucinated facts. The research suggests that current AI models (at least the Llama-3 8B model) have yet to be maturely developed enough to serve as reliable tools for global technical governance. The digital barriers and gaps identified in this paper must be addressed through more inclusive data representation in model training and more transparent alignment processes to ensure that AI benefits—including safety, fairness and readiness—are accessible to all countries, regardless of their geographical location or income classification.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The research question is important and well-motivated. If policymakers increasingly rely on LLMs, geographic knowledge gaps have real governance consequences, and the GAID dataset provides a solid ground-truth benchmark. However, the experimental design conflates the model's training data limitations with bias. The author acknowledges that Llama-3 8B may not have access to 2025 data, but still draws strong conclusions about geographic exclusion without controlling for this confound.

  2. The authors propose a reasonable strategy for measuring a policy-relevant model capability—ensuring models are not substantively worse at answering questions about some countries versus others. My core criticism is that they do not spend enough time thinking carefully about what information their evaluation strategy actually captures. The query year postdates the model's training cutoff, so refusals may reflect temporal limitations rather than geographic bias. And the response categories are too coarse. More time spent reading through outputs and interpreting what different failure modes actually mean would strengthen the conclusions considerably.

Cite this project

@misc{hung2026global,
  title = {{Global AI Bias Audit for Technical Governance}},
  author = {Jason Hung},
  year = {2026},
  month = feb,
  note = {Submitted to The Technical AI Governance Challenge, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/global-ai-bias-audit-for-technical-governance-q1t8}},
  url = {https://apartresearch.com/sprints/projects/global-ai-bias-audit-for-technical-governance-q1t8}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026