Cross-Lingual Safety Audit of LLMs in South African Languages
Thando Shabangu
We test whether four LLMs (Claude Sonnet 4.6 plus three locally-run open-weight models) keep their safety guardrails when harmful requests are issued in three South African languages — isiZulu, Tshivenḓa, Sepedi — versus English. The open-weight models refuse 69% of harmful prompts in English but only 7% in the indigenous languages, and we show this degradation splits into genuine safety failures and mere capability failures — a distinction we argue is essential for honest multilingual safety evaluation.
A clean, well-represented audit. the 4-way coding scheme that separates safety failures from capability failures is interesting. That distinction matters and is underappreciated in the literature. The small sample size is the main limitation
Strengths:
- The 4-way response coding (refusal / partial / full compliance / alignment failure) is a real contribution. When a model produces gibberish in Tshivenḓa, that's not a safety failure, it's a capability failure. Most other work conflates these.
- Human annotation rather than LLM-as-judge adds credibility, especially for low-resource languages.
Suggestions for Future Work:
- With only 12 prompts per cell, the quantitative claims are on shaky ground. Scaling this up would make the findings much more convincing.
- Including 4B-parameter models that obviously cannot process the target languages inflates the apparent safety gap. Matching model capabilities to the languages being tested would give a cleaner picture.
-- Impact potential & innovation
++ execution quality
++ presentation & clarity
The project aims to characterize how different LLM’s respond to harmful request in native South African languages such as isiZulu, Tshivenda, and Sepedi. The project identifies an important limitation of LLMs to handle malicious requests in these native languages, particularly in open models such as Qwen and Gemma. It is also highlighted the distinction between compliance and capability failures of the LLMs to understand these languages. Nevertheless, the sample size of the experiments is not large enough to draw enough conclusions.
Cite this work
@misc {
title={
(HckPrj) Cross-Lingual Safety Audit of LLMs in South African Languages
},
author={
Thando Shabangu
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


