Skip to content
Sprint projectJun 22, 2026Nairobi, Kenya

AfriSafe-CB: Evaluating LLM Safety Robustness Under African Code Switched Political and Civic Contexts

Michelle Wanjiku Thuo, Mahmoud Mannes · Team AfriGuard AI

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: AfriSafe-CB: Evaluating LLM Safety Robustness Under African Code Switched Political and Civic Contexts

Share

Artificial intelligence is becoming part of how people learn, access information, make decisions, and participate in society. However, most AI safety testing is still designed around English conversations, leaving an important question unanswered: do AI systems remain reliable when people communicate in the multilingual and code switched ways that are common across Africa?

AfriSafe-CB (African Code-Switched Safety Benchmark) is a benchmark designed to explore this gap. It tests whether large language models can maintain safe, accurate, and responsible behaviour when faced with safety sensitive situations expressed through different language contexts, including English, Sheng/Swahili code switching, and Arabic/French code switching.

The benchmark contains 50 carefully designed scenarios covering real world challenges such as misinformation, election-related claims, phishing attempts, deepfakes, fraud, online manipulation, and information integrity. Each scenario is evaluated across different language conditions to identify whether AI systems understand context, resist harmful requests, and provide reliable guidance consistently.

By focusing on communication patterns often overlooked in AI evaluation, AfriSafe-CB aims to contribute toward building AI systems that are safer, fairer, and more trustworthy for diverse communities. The project provides an initial framework for researchers, developers, and policymakers to better understand how AI safety performs beyond English centric environments.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. I would reject this.

    The primary reason for rejection is that this submission reads as a research protocol or proposal rather than a completed study — it describes expected findings and expected visualizations but reports zero empirical results, with the example scores presented being illustrative rather than actual data. This alone is disqualifying for any venue expecting completed research. Beyond this fundamental issue, the project enters a space that has become significantly crowded in 2025–2026, with multiple completed papers already answering essentially the same research question with actual data at larger scales, including the LSR Benchmark for cross-lingual refusal degradation in West African languages, a study on multilingual jailbreaking via low-resource African languages including Kiswahili, a culturally-grounded policy benchmark for equitable AI safety in African languages, and RabakBench which evaluated 13 guardrails with rigorous human-in-the-loop validation achieving 0.70–0.80 inter-annotator agreement. The benchmark itself is too small at only 50 prompts across 3 linguistic conditions, which is insufficient to draw statistically meaningful conclusions, especially when competing work operates at significantly larger scales. The code-switching angle, while interesting, does not sufficiently differentiate from existing work that already tests Kiswahili and other African languages in safety contexts. Methodologically, the paper lacks inter-annotator agreement protocols, statistical significance testing, and confidence intervals, and it acknowledges that human annotation may introduce subjectivity without proposing a solution. Finally, testing only three models (ChatGPT, DeepSeek, Gemini) is too narrow for 2026 when the field has moved to evaluating thirteen or more guardrail systems, and the literature review does not engage with or differentiate from the substantial body of closely related work published in the past year.

    Read full reviewShow less
  2. There is real thought in this design. Going after code-switching specifically, Sheng mixed with Swahili, Arabic mixed with French, is closer to how people actually talk than translating a whole prompt into one language, and the civic and election angle is timely. Your failure categories are sharp too, separating "the model misunderstood" from "the model complied unsafely" is exactly the right line to draw. The workbook and annotation guide are a real, usable artifact. The one thing missing is the part that turns this into a result: you have not run it yet. You have the prompts, the three language conditions, and the scoring all set, so running even two or three of the models you listed across the 50 prompts would give you real numbers and let you fill in those figures. That is basically a weekend of API calls, and you are closer than it probably feels. Good foundation to build on.

  3. The project aims to produce a safety benchmark to analyze how different LLMs react to malicious requests across different types of misinformation. Specifically, the authors want to examine this in different languages and their switch variants (eg. Sheng/Swahili). Although the intention of the project is good there are little results included for proper evaluation to give a verdict for the models it tries to tackle.

Cite this project

@misc{thuo2026afrisafecb,
  title = {{AfriSafe-CB: Evaluating LLM Safety Robustness Under African Code Switched Political and Civic Contexts}},
  author = {Michelle Wanjiku Thuo and Mahmoud Mannes},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/afrisafecb-evaluating-llm-safety-robustness-under-african-code-switched-political-and-civic-contexts-x8tn}},
  url = {https://apartresearch.com/sprints/projects/afrisafecb-evaluating-llm-safety-robustness-under-african-code-switched-political-and-civic-contexts-x8tn}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026