Skip to content
Sprint projectJun 21, 2026Abuja

Multilingual Jailbreak Vulnerability Benchmark and Mitigation for Low-Resource African Languages

Godwin Abuh Faruna · Team Godwin Abuh Faruna (solo)

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Multilingual Jailbreak Vulnerability Benchmark and Mitigation for Low-Resource African Languages

Code (opens in new tab)More on huggingface.co (opens in new tab)
Share

We present a two-part study: (1) a multilingual jailbreak benchmark revealing large safety gaps in four open-weight LLMs across seven languages including Igala, which has no prior AI safety coverage, and (2) Latent Space Refusal Anchoring (LSR-Anchoring), a training-free activation-steering mitigation that recovers safety at inference time without retraining or target-language data. The combined result is a full pipeline: we find the failure, characterise it geometrically, and fix it, all without supervised African-language data. Across four models tested, English SRR stays at 85–99%, while African language SRR collapses to single digits in some cases. LSR-Anchoring recovers safety to near-ceiling on four languages at 70B scale, with MMLU accuracy drops below 0.35 percentage points. Arabic fails on all methods due to geometric misalignment, providing a clear deployment decision rule.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. Very well executed. It goes beyond benchmarking to actually propose and implement a fix (training-free activation steering to recover cross-lingual safety). The work is ambitious and mostly delivers.

    Strengths:

    - Actually builds a mitigation, not just a benchmark. The LSR-Anchoring method works at inference time without retraining, which is practically useful.

    - The Arabic negative result (English-derived steering makes things worse) is genuinely interesting and reveals something about the geometry of the refusal manifold.

    Areas for improvements

    - basically the limitations mentioned in the paper. Using a semantic classifier instead of phrase-based, multi-human review, using obfuscated prompts, etc

    - slides missing

  2. Great project. Can be expanded to solve multiple other problems within the space. Innovative and good execution quality but can be presented better.

  3. The project evaluates different open source LLMs with different African languages searching for jailbreak vulnerabilities that are possible due to the model’s overfitting in English language information. The authors successfully the problem LLM’s have with these languages and offer a solution by using Latent Space Refusal anchoring to add the English refusal direction to the model’s response to counter the jailbreak while also considering that the jailbreak could be a consequence of model incapability to understand the language. However, the way that the mitigation approach works is not sufficiently explained in the document with missing terms in crucial equations. Likewise, the paper could use grounded examples of some prompts and responses with and without the jailbreak (while blurring sensitive data) to demonstrate how the proposed defense works practically.

Cite this project

@misc{faruna2026multilingual,
  title = {{Multilingual Jailbreak Vulnerability Benchmark and Mitigation for Low-Resource African Languages}},
  author = {Godwin Abuh Faruna},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/multilingual-jailbreak-vulnerability-benchmark-and-mitigation-for-lowresource-african-languages-td0s}},
  url = {https://apartresearch.com/sprints/projects/multilingual-jailbreak-vulnerability-benchmark-and-mitigation-for-lowresource-african-languages-td0s}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026