Cross-Lingual Safety Audit of LLMs in South African Languages
Thando Shabangu · Team Tech Titans
Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
We test whether four LLMs (Claude Sonnet 4.6 plus three locally-run open-weight models) keep their safety guardrails when harmful requests are issued in three South African languages — isiZulu, Tshivenḓa, Sepedi — versus English. The open-weight models refuse 69% of harmful prompts in English but only 7% in the indigenous languages, and we show this degradation splits into genuine safety failures and mere capability failures — a distinction we argue is essential for honest multilingual safety evaluation.
Reviews
A clean, well-represented audit. the 4-way coding scheme that separates safety failures from capability failures is interesting. That distinction matters and is underappreciated in the literature. The small sample size is the main limitation
Strengths:
- The 4-way response coding (refusal / partial / full compliance / alignment failure) is a real contribution. When a model produces gibberish in Tshivenḓa, that's not a safety failure, it's a capability failure. Most other work conflates these.
- Human annotation rather than LLM-as-judge adds credibility, especially for low-resource languages.
Suggestions for Future Work:
- With only 12 prompts per cell, the quantitative claims are on shaky ground. Scaling this up would make the findings much more convincing.
- Including 4B-parameter models that obviously cannot process the target languages inflates the apparent safety gap. Matching model capabilities to the languages being tested would give a cleaner picture.
Read full reviewShow less
-- Impact potential & innovation
++ execution quality
++ presentation & clarity
The project aims to characterize how different LLM’s respond to harmful request in native South African languages such as isiZulu, Tshivenda, and Sepedi. The project identifies an important limitation of LLMs to handle malicious requests in these native languages, particularly in open models such as Qwen and Gemma. It is also highlighted the distinction between compliance and capability failures of the LLMs to understand these languages. Nevertheless, the sample size of the experiments is not large enough to draw enough conclusions.
Cite this project
@misc{shabangu2026crosslingual,
title = {{Cross-Lingual Safety Audit of LLMs in South African Languages}},
author = {Thando Shabangu},
year = {2026},
month = jun,
note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/crosslingual-safety-audit-of-llms-in-south-african-languages-dtp3}},
url = {https://apartresearch.com/sprints/projects/crosslingual-safety-audit-of-llms-in-south-african-languages-dtp3}
}More from Global South AI Safety Hackathon
- View project: Pragmatic Sophistry in Vietnamese Multi-Agent Oversight
Pragmatic Sophistry in Vietnamese Multi-Agent Oversight
AI Safety Enthusiasts
AI safety monitors are usually evaluated on the assumption that risky behavior is lexically visible in the text being watched. We test this assumption in a multilingual, multi-agent setting: Vietnamese-language workflow …
- View project: JusticIA: A Counterfactual Benchmark for Auditing Contextual Biases in Language Models for Transitional Justice
JusticIA: A Counterfactual Benchmark for Auditing Contextual Biases in Language Models for Transitional Justice
JusticeMiners
JusticIA is a counterfactual benchmark for auditing contextual bias in LLMs applied to Colombian transitional justice. It tests whether six LLMs change their sanction recommendations when only one contextual attribute …
- View project: Coldron
Coldron
ColDron
En Colombia, los grupos armados ilegales ya atacan con drones comerciales modificados y ya han herido y matado a civiles. Una pregunta decide cómo gobernar esta amenaza: ¿quién elige el blanco y aprieta el gatillo? Hoy, …