Skip to content
Sprint projectJun 22, 2026Mexico

Cultural Knowledge Gaps in LLMs: Geographic Hallucination Bias Across Latin American Countries

Karla Angelica Doctor Mauricio · Team Fairness LATAM

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Cultural Knowledge Gaps in LLMs: Geographic Hallucination Bias Across Latin American Countries

Code (opens in new tab)
Share

Large language models (LLMs) are used in education, public services, and information tools across Latin America. But we do not know how often they produce wrong information about Latin American culture, or where they fail the most. This paper evaluates two models, GPT-4o-mini and Claude Haiku 4.5, on 270 questions from CHOCLO, a benchmark of cultural knowledge from 18 Latin American countries. We use semantic similarity scores and a four-category LLM-as-judge classifier (correct, partial, hallucination, abstention) to compare how each model fails. GPT-4o-mini hallucinates at 33.0% overall and almost never abstains (1.5%). Claude Haiku 4.5 hallucinates less (24.1%) but abstains much more (17.8%). Both models show geographic gaps of 33 to 40 percentage points across countries. These gaps remain after controlling for category composition and topic diversity. The public_figure category has the highest hallucination rate (66.7% for GPT), and it is also the least represented category in the benchmark itself. This points to a two-layer problem: models lack knowledge about specific Latin American individuals, and the benchmark has limited coverage of them too. We also document infrastructure barriers that prevented evaluation of LatamGPT. We propose three data governance models as a path forward for communities in the region.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. - Very well written and organized. The research question is very clearly stated.

    - Very good methodology.

    - I really liked the limitations section, being conscious of the lack of indigenous and other population representation, and the proposed next steps.

    - To verify that the gap is genuinely specific to Latin America and not just a general hallucination issue, it would be valuable to include a comparison with: (1) the same questions in English (to isolate the effect of language), and (2) questions specific to the US or the Global North (to have an actual comparison point).

    - It would be good to see an analysis of abstention itself. What kinds of questions does Claude, for example, avoid answering? Are they questions related to, for instance, armed conflict in the region?

    - I noticed that CHOCLO (on Hugging Face) already runs an evaluation of its dataset against several models. The paper doesn't make clear what the actual innovation is relative to that. It might be the abstention/failure-mode analysis, but the paper should state this much more explicitly.

    Read full reviewShow less
  2. The underlying question is well chosen: what happens when a chatbot fails with confidence on a cultural topic from Guatemala or Venezuela, rather than admitting it doesn't know, and the contrast between GPT-4o-mini hallucinating fluently and Claude abstaining connects well with the safety argument that actually matters, which is that the silent failure is the dangerous one. That said, the methodological novelty is more limited than it first appears: the team itself describes its contribution as an adaptation of existing fairness tools, bootstrap, UMAP, multiple comparison corrections, to a new domain, rather than a method of its own, and even the governance proposals are explicitly inspired by Masakhane.

    What weighs most heavily, though, is execution. The team flags that GPT-4o-mini judges its own responses and Claude's, but never validates whether that's actually biasing the results , no human check, just a noted limitation. The repository makes this worse: the hallucination-by-category numbers in the README don't match the ones in the paper. GPT's geography hallucination goes from 25.5% to 37.3%. Claude's public_figure goes from 22.2% to 33.3%. Nothing explains which run the published conclusions actually rest on. That's not a presentation slip, it's a validation gap in the study's own numbers.

    On presentation, the title speaks of Latin America broadly when Brazil and Haiti are excluded due to CHOCLO's own limits, and Table 4 mixes very wide confidence intervals, such as Argentina's 26.7-80.0%, with fairly firm claims of robustness in the surrounding text, without flagging that with n=15 per country this is to be expected.

    Overall, a solid and well-framed impact question, built on standard tools rather than a method of its own, with a validation gap that outweighs the presentation details.

    Read full reviewShow less
  3. This project produces directly actionable empirical evidence about a real AI safety risk in LATAM. The qualitative examples in Table 2 are particularly effective for communicating the problem to non-technical audiences. To strengthen the work, the most urgent next steps are increasing to at least 30 questions per country to reduce confidence intervals, separating the judge role from the evaluated model (using a third independent model as judge), and expanding coverage of the public_figure category, which is precisely where hallucination rates are highest. The Representativeness Index proposal and the Masakhane-inspired governance models are concrete, well-grounded policy contributions that deserve further development.

Cite this project

@misc{mauricio2026cultural,
  title = {{Cultural Knowledge Gaps in LLMs: Geographic Hallucination Bias Across Latin American Countries}},
  author = {Karla Angelica Doctor Mauricio},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/cultural-knowledge-gaps-in-llms-geographic-hallucination-bias-across-latin-american-countries-5px9}},
  url = {https://apartresearch.com/sprints/projects/cultural-knowledge-gaps-in-llms-geographic-hallucination-bias-across-latin-american-countries-5px9}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026