Skip to content
Sprint projectJun 21, 2026Johannesburg

AfriSafe-Eval

Tebogo Jan Seopa · Team DevRift

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

AfriSafe-Eval is a 400-prompt red-teaming benchmark testing LLM safety across five South African languages and four locally-grounded harm categories: electoral manipulation, healthcare misinformation, financial fraud, and GBV facilitation. Across 1,600 responses from four LLMs, harmful response rates ranged from 17.8% to 61.2% by language. isiXhosa was riskiest in every model tested, isiZulu among the safest, despite comparable resourcing, showing the safety gap isn't simply about data scarcity. Dataset, harness, and validation pipeline are fully open source.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Tebogo — the core finding here is worth taking seriously. Writing the prompts natively instead of translating them, and then seeing a ~37-point isiXhosa/isiZulu gap survive across four models, is the kind of result that's genuinely hard to dismiss. The honest weak point is the measurement layer underneath it.

    Your harmful-compliance numbers are only as trustworthy as the classifier producing them, and that classifier was validated on 40 responses with what looks like a single annotator and no per-language breakdown. So the first thing I'd want is a second labeler on a bigger sample and a Cohen's kappa (or Gwet's AC1, given the class imbalance) — without an agreement number, reviewers can't tell signal from labeling noise. Report per-class specificity and the confusion matrix too, not just aggregate accuracy; a rule-based labeler that only reads the first ~250 characters will miss late refusals and corrective framing, and there's no reason to assume it misses them evenly across five languages.

    Two smaller things. You sampled each model once, so there are no confidence intervals on any rate. And the four-model set was picked off free Huawei credits — worth stating plainly so nobody reads it as a designed comparison.

    Read full reviewShow less
  2. The benchmark created is solid and the results - differences with English - are striking. I have a concern about the main finding - that cross lingual safety gaps in LLMs are not well explained by resource availability on its own. Is representation in CC a great proxy for the dataset frontier models are trained on? Both languages have such low representation in CC that adding even small amounts of (ideally high quality) data for either language from any other source can greatly affect total representation. Would be more conclusive if this could be run on open source models that make all their training data public (Pythia? OLMo 2?) + some analysis on the training data to support this.

  3. Interesting findings on divergence in harmful response rate between isiXhosa (61.2%) and isiZulu (23.8%)!

    One reason this could happen is if there are differences in prompts introduced through translation: isiXhosa-vs-isiZulu could differ because the isiXhosa prompts are simply more concrete, more fluent, or push harder, not because the models are less safe in isiXhosa.

    Consider doing control experiments (e.g. having a human rate prompts based on difficulty or harm) to isolate whether there is authorship variance.

  4. Well done. I especially liked that the prompts were natively written rather than translated, and that the human labeled validation sample was included, which adds credibility compared with fully automated evaluations. The isiXhosa versus isiZulu finding is particularly interesting and novel.

Cite this project

@misc{seopa2026afrisafeeval,
  title = {{AfriSafe-Eval}},
  author = {Tebogo Jan Seopa},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/afrisafeeval-mc07}},
  url = {https://apartresearch.com/sprints/projects/afrisafeeval-mc07}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026