AfriSafe-Eval
Tebogo Jan Seopa · Team DevRift
Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
AfriSafe-Eval is a 400-prompt red-teaming benchmark testing LLM safety across five South African languages and four locally-grounded harm categories: electoral manipulation, healthcare misinformation, financial fraud, and GBV facilitation. Across 1,600 responses from four LLMs, harmful response rates ranged from 17.8% to 61.2% by language. isiXhosa was riskiest in every model tested, isiZulu among the safest, despite comparable resourcing, showing the safety gap isn't simply about data scarcity. Dataset, harness, and validation pipeline are fully open source.

Reviews
Tebogo — the core finding here is worth taking seriously. Writing the prompts natively instead of translating them, and then seeing a ~37-point isiXhosa/isiZulu gap survive across four models, is the kind of result that's genuinely hard to dismiss. The honest weak point is the measurement layer underneath it.
Your harmful-compliance numbers are only as trustworthy as the classifier producing them, and that classifier was validated on 40 responses with what looks like a single annotator and no per-language breakdown. So the first thing I'd want is a second labeler on a bigger sample and a Cohen's kappa (or Gwet's AC1, given the class imbalance) — without an agreement number, reviewers can't tell signal from labeling noise. Report per-class specificity and the confusion matrix too, not just aggregate accuracy; a rule-based labeler that only reads the first ~250 characters will miss late refusals and corrective framing, and there's no reason to assume it misses them evenly across five languages.
Two smaller things. You sampled each model once, so there are no confidence intervals on any rate. And the four-model set was picked off free Huawei credits — worth stating plainly so nobody reads it as a designed comparison.
Read full reviewShow less
The benchmark created is solid and the results - differences with English - are striking. I have a concern about the main finding - that cross lingual safety gaps in LLMs are not well explained by resource availability on its own. Is representation in CC a great proxy for the dataset frontier models are trained on? Both languages have such low representation in CC that adding even small amounts of (ideally high quality) data for either language from any other source can greatly affect total representation. Would be more conclusive if this could be run on open source models that make all their training data public (Pythia? OLMo 2?) + some analysis on the training data to support this.
Interesting findings on divergence in harmful response rate between isiXhosa (61.2%) and isiZulu (23.8%)!
One reason this could happen is if there are differences in prompts introduced through translation: isiXhosa-vs-isiZulu could differ because the isiXhosa prompts are simply more concrete, more fluent, or push harder, not because the models are less safe in isiXhosa.
Consider doing control experiments (e.g. having a human rate prompts based on difficulty or harm) to isolate whether there is authorship variance.
Well done. I especially liked that the prompts were natively written rather than translated, and that the human labeled validation sample was included, which adds credibility compared with fully automated evaluations. The isiXhosa versus isiZulu finding is particularly interesting and novel.
Cite this project
@misc{seopa2026afrisafeeval,
title = {{AfriSafe-Eval}},
author = {Tebogo Jan Seopa},
year = {2026},
month = jun,
note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/afrisafeeval-mc07}},
url = {https://apartresearch.com/sprints/projects/afrisafeeval-mc07}
}More from Global South AI Safety Hackathon
- View project: Pragmatic Sophistry in Vietnamese Multi-Agent Oversight
Pragmatic Sophistry in Vietnamese Multi-Agent Oversight
AI Safety Enthusiasts
AI safety monitors are usually evaluated on the assumption that risky behavior is lexically visible in the text being watched. We test this assumption in a multilingual, multi-agent setting: Vietnamese-language workflow …
- View project: JusticIA: A Counterfactual Benchmark for Auditing Contextual Biases in Language Models for Transitional Justice
JusticIA: A Counterfactual Benchmark for Auditing Contextual Biases in Language Models for Transitional Justice
JusticeMiners
JusticIA is a counterfactual benchmark for auditing contextual bias in LLMs applied to Colombian transitional justice. It tests whether six LLMs change their sanction recommendations when only one contextual attribute …
- View project: Coldron
Coldron
ColDron
En Colombia, los grupos armados ilegales ya atacan con drones comerciales modificados y ya han herido y matado a civiles. Una pregunta decide cómo gobernar esta amenaza: ¿quién elige el blanco y aprieta el gatillo? Hoy, …