Cultural Knowledge Gaps in LLMs: Geographic Hallucination Bias Across Latin American Countries
Karla Angelica Doctor Mauricio
Large language models (LLMs) are used in education, public services, and information tools across Latin America. But we do not know how often they produce wrong information about Latin American culture, or where they fail the most. This paper evaluates two models, GPT-4o-mini and Claude Haiku 4.5, on 270 questions from CHOCLO, a benchmark of cultural knowledge from 18 Latin American countries. We use semantic similarity scores and a four-category LLM-as-judge classifier (correct, partial, hallucination, abstention) to compare how each model fails. GPT-4o-mini hallucinates at 33.0% overall and almost never abstains (1.5%). Claude Haiku 4.5 hallucinates less (24.1%) but abstains much more (17.8%). Both models show geographic gaps of 33 to 40 percentage points across countries. These gaps remain after controlling for category composition and topic diversity. The public_figure category has the highest hallucination rate (66.7% for GPT), and it is also the least represented category in the benchmark itself. This points to a two-layer problem: models lack knowledge about specific Latin American individuals, and the benchmark has limited coverage of them too. We also document infrastructure barriers that prevented evaluation of LatamGPT. We propose three data governance models as a path forward for communities in the region.
- Very well written and organized. The research question is very clearly stated.
- Very good methodology.
- I really liked the limitations section, being conscious of the lack of indigenous and other population representation, and the proposed next steps.
- To verify that the gap is genuinely specific to Latin America and not just a general hallucination issue, it would be valuable to include a comparison with: (1) the same questions in English (to isolate the effect of language), and (2) questions specific to the US or the Global North (to have an actual comparison point).
- It would be good to see an analysis of abstention itself. What kinds of questions does Claude, for example, avoid answering? Are they questions related to, for instance, armed conflict in the region?
- I noticed that CHOCLO (on Hugging Face) already runs an evaluation of its dataset against several models. The paper doesn't make clear what the actual innovation is relative to that. It might be the abstention/failure-mode analysis, but the paper should state this much more explicitly.
The underlying question is well chosen: what happens when a chatbot fails with confidence on a cultural topic from Guatemala or Venezuela, rather than admitting it doesn't know, and the contrast between GPT-4o-mini hallucinating fluently and Claude abstaining connects well with the safety argument that actually matters, which is that the silent failure is the dangerous one. That said, the methodological novelty is more limited than it first appears: the team itself describes its contribution as an adaptation of existing fairness tools, bootstrap, UMAP, multiple comparison corrections, to a new domain, rather than a method of its own, and even the governance proposals are explicitly inspired by Masakhane.
What weighs most heavily, though, is execution. The team flags that GPT-4o-mini judges its own responses and Claude's, but never validates whether that's actually biasing the results , no human check, just a noted limitation. The repository makes this worse: the hallucination-by-category numbers in the README don't match the ones in the paper. GPT's geography hallucination goes from 25.5% to 37.3%. Claude's public_figure goes from 22.2% to 33.3%. Nothing explains which run the published conclusions actually rest on. That's not a presentation slip, it's a validation gap in the study's own numbers.
On presentation, the title speaks of Latin America broadly when Brazil and Haiti are excluded due to CHOCLO's own limits, and Table 4 mixes very wide confidence intervals, such as Argentina's 26.7-80.0%, with fairly firm claims of robustness in the surrounding text, without flagging that with n=15 per country this is to be expected.
Overall, a solid and well-framed impact question, built on standard tools rather than a method of its own, with a validation gap that outweighs the presentation details.
This project produces directly actionable empirical evidence about a real AI safety risk in LATAM. The qualitative examples in Table 2 are particularly effective for communicating the problem to non-technical audiences. To strengthen the work, the most urgent next steps are increasing to at least 30 questions per country to reduce confidence intervals, separating the judge role from the evaluated model (using a third independent model as judge), and expanding coverage of the public_figure category, which is precisely where hallucination rates are highest. The Representativeness Index proposal and the Masakhane-inspired governance models are concrete, well-grounded policy contributions that deserve further development.
Cite this work
@misc {
title={
(HckPrj) Cultural Knowledge Gaps in LLMs: Geographic Hallucination Bias Across Latin American Countries
},
author={
Karla Angelica Doctor Mauricio
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


