Global AI Bias Audit for Technical Governance
Jason Hung · Team Global AI Dataset Project
Submitted to The Technical AI Governance Challenge. Sprint projects are early-stage work by participants, not Apart Research publications.
This project is the exploratory phase of Phases 3-4 of my milestone-based, ongoing Global AI Dataset (GAID) Project. In this exploratory project, I used the version 2 GAID dataset (published on Harvard Dataverse) as a framework to stress-test the open-weight Llama-3 8B model and evaluate geographic and socioeconomic biases in technical AI governance awareness. By stress-testing the model with 1,704 queries across 213 countries and eight technical metrics, I identified a significant digital barrier and gap separating the Global North and South. The results indicate that the model was only able to provide number/fact responses in 11.4% of its query answers, where the empirical validity of such responses was yet to be verified. The findings reveal that AI's technical knowledge is heavily concentrated in higher-income regions, while lower-income countries from the Global South are subject to disproportionate systemic information gaps. This disparity between the Global North and South poses concerning risks for global AI safety and inclusive governance, as policymakers in underserved regions may lack reliable data-driven insights or be misled by hallucinated facts. The research suggests that current AI models (at least the Llama-3 8B model) have yet to be maturely developed enough to serve as reliable tools for global technical governance. The digital barriers and gaps identified in this paper must be addressed through more inclusive data representation in model training and more transparent alignment processes to ensure that AI benefits—including safety, fairness and readiness—are accessible to all countries, regardless of their geographical location or income classification.
Reviews
The research question is important and well-motivated. If policymakers increasingly rely on LLMs, geographic knowledge gaps have real governance consequences, and the GAID dataset provides a solid ground-truth benchmark. However, the experimental design conflates the model's training data limitations with bias. The author acknowledges that Llama-3 8B may not have access to 2025 data, but still draws strong conclusions about geographic exclusion without controlling for this confound.
The authors propose a reasonable strategy for measuring a policy-relevant model capability—ensuring models are not substantively worse at answering questions about some countries versus others. My core criticism is that they do not spend enough time thinking carefully about what information their evaluation strategy actually captures. The query year postdates the model's training cutoff, so refusals may reflect temporal limitations rather than geographic bias. And the response categories are too coarse. More time spent reading through outputs and interpreting what different failure modes actually mean would strengthen the conclusions considerably.
Cite this project
@misc{hung2026global,
title = {{Global AI Bias Audit for Technical Governance}},
author = {Jason Hung},
year = {2026},
month = feb,
note = {Submitted to The Technical AI Governance Challenge, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/global-ai-bias-audit-for-technical-governance-q1t8}},
url = {https://apartresearch.com/sprints/projects/global-ai-bias-audit-for-technical-governance-q1t8}
}More from The Technical AI Governance Challenge
- 1st placeView project: LidaSim: Testing AI Policies With Persona-Based Simulations
LidaSim: Testing AI Policies With Persona-Based Simulations
Lida Safety
We simulate well-known figures in AI and politics with agents, scraping large amounts of data to get realistic simulations. Then, we test questions and proposed policies against these public figures, to see which …
- 2nd placeView project: Markov Chain Lock Watermarking: Provably Secure Authentication for LLM Outputs
Markov Chain Lock Watermarking: Provably Secure Authentication for LLM Outputs
MCL
We present Markov Chain Lock (MCL) watermarking, a cryptographically secure framework for authenticating LLM outputs. MCL constrains token generation to follow a secret Markov chain over SHA-256 vocabulary partitions. …
- 3rd placeView project: Political Intelligence for AI Safety: The AI Risk Attitudes Survey (AIRAS)
Political Intelligence for AI Safety: The AI Risk Attitudes Survey (AIRAS)
AIRAS
The AI safety and governance community is making progress on defining red lines around existential risk from advanced AI systems, and building verification infrastructure to support this objective. However, this is only …