LLMs Flatten the Global South: Sub-Regional Representation Asymmetry
Noah De Nicola · Team Noah De Nicola (solo)
Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
LLMs default to Northern values and collapse within-country variation. Prior work shows this at the output level. We ask whether the same asymmetry shows up in the model's internal representations, with sub-regions in the Global South separating less than those in the Global North. We build persona vectors for 21 countries and 105 cities across three open-weight LLMs, and measure how far each country's cities separate from it in activation space. We control for two confounds: Wikipedia corpus frequency and BPE tokenizer fragmentation. Global North sub-regions separate 39% more than Global South across all layers and 77% more in later layers. The gap survives matched corpus frequency and token fragmentation.
Reviews
I really liked the angle of this paper. It does not just look at what the model says, but tries to understand what the model actually represents internally. That makes the work feel more original and deeper than a typical bias evaluation.
The most interesting insight for me is that Global South regions may be getting flattened inside the model itself, not just in the final answers. That is an important point because it suggests cultural fairness cannot be solved only by better prompting or nicer wording. I would also be interested in seeing this extended to countries with much higher internal diversity, such as India, where regional identity is shaped not only by cities but also by language, caste, religion, ethnicity, and local community. That kind of setting could make the paper’s core question even more powerful: whether models truly represent internal cultural diversity, or flatten complex societies into one broad national identity.
Read full reviewShow less
4 / 4 / 4
Good and important topic. The method is solid, but it needs stronger checks to prove the difference is really cultural representation and not another hidden factor.
Cite this project
@misc{nicola2026llms,
title = {{LLMs Flatten the Global South: Sub-Regional Representation Asymmetry}},
author = {Noah De Nicola},
year = {2026},
month = jun,
note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/llms-flatten-the-global-south-subregional-representation-asymmetry-mw2b}},
url = {https://apartresearch.com/sprints/projects/llms-flatten-the-global-south-subregional-representation-asymmetry-mw2b}
}More from Global South AI Safety Hackathon
- View project: Pragmatic Sophistry in Vietnamese Multi-Agent Oversight
Pragmatic Sophistry in Vietnamese Multi-Agent Oversight
AI Safety Enthusiasts
AI safety monitors are usually evaluated on the assumption that risky behavior is lexically visible in the text being watched. We test this assumption in a multilingual, multi-agent setting: Vietnamese-language workflow …
- View project: JusticIA: A Counterfactual Benchmark for Auditing Contextual Biases in Language Models for Transitional Justice
JusticIA: A Counterfactual Benchmark for Auditing Contextual Biases in Language Models for Transitional Justice
JusticeMiners
JusticIA is a counterfactual benchmark for auditing contextual bias in LLMs applied to Colombian transitional justice. It tests whether six LLMs change their sanction recommendations when only one contextual attribute …
- View project: Coldron
Coldron
ColDron
En Colombia, los grupos armados ilegales ya atacan con drones comerciales modificados y ya han herido y matado a civiles. Una pregunta decide cómo gobernar esta amenaza: ¿quién elige el blanco y aprieta el gatillo? Hoy, …