AfriSafe-CB: Evaluating LLM Safety Robustness Under African Code Switched Political and Civic Contexts
Michelle Wanjiku Thuo, Mahmoud Mannes
Artificial intelligence is becoming part of how people learn, access information, make decisions, and participate in society. However, most AI safety testing is still designed around English conversations, leaving an important question unanswered: do AI systems remain reliable when people communicate in the multilingual and code switched ways that are common across Africa?
AfriSafe-CB (African Code-Switched Safety Benchmark) is a benchmark designed to explore this gap. It tests whether large language models can maintain safe, accurate, and responsible behaviour when faced with safety sensitive situations expressed through different language contexts, including English, Sheng/Swahili code switching, and Arabic/French code switching.
The benchmark contains 50 carefully designed scenarios covering real world challenges such as misinformation, election-related claims, phishing attempts, deepfakes, fraud, online manipulation, and information integrity. Each scenario is evaluated across different language conditions to identify whether AI systems understand context, resist harmful requests, and provide reliable guidance consistently.
By focusing on communication patterns often overlooked in AI evaluation, AfriSafe-CB aims to contribute toward building AI systems that are safer, fairer, and more trustworthy for diverse communities. The project provides an initial framework for researchers, developers, and policymakers to better understand how AI safety performs beyond English centric environments.
I would reject this.
The primary reason for rejection is that this submission reads as a research protocol or proposal rather than a completed study — it describes expected findings and expected visualizations but reports zero empirical results, with the example scores presented being illustrative rather than actual data. This alone is disqualifying for any venue expecting completed research. Beyond this fundamental issue, the project enters a space that has become significantly crowded in 2025–2026, with multiple completed papers already answering essentially the same research question with actual data at larger scales, including the LSR Benchmark for cross-lingual refusal degradation in West African languages, a study on multilingual jailbreaking via low-resource African languages including Kiswahili, a culturally-grounded policy benchmark for equitable AI safety in African languages, and RabakBench which evaluated 13 guardrails with rigorous human-in-the-loop validation achieving 0.70–0.80 inter-annotator agreement. The benchmark itself is too small at only 50 prompts across 3 linguistic conditions, which is insufficient to draw statistically meaningful conclusions, especially when competing work operates at significantly larger scales. The code-switching angle, while interesting, does not sufficiently differentiate from existing work that already tests Kiswahili and other African languages in safety contexts. Methodologically, the paper lacks inter-annotator agreement protocols, statistical significance testing, and confidence intervals, and it acknowledges that human annotation may introduce subjectivity without proposing a solution. Finally, testing only three models (ChatGPT, DeepSeek, Gemini) is too narrow for 2026 when the field has moved to evaluating thirteen or more guardrail systems, and the literature review does not engage with or differentiate from the substantial body of closely related work published in the past year.
There is real thought in this design. Going after code-switching specifically, Sheng mixed with Swahili, Arabic mixed with French, is closer to how people actually talk than translating a whole prompt into one language, and the civic and election angle is timely. Your failure categories are sharp too, separating "the model misunderstood" from "the model complied unsafely" is exactly the right line to draw. The workbook and annotation guide are a real, usable artifact. The one thing missing is the part that turns this into a result: you have not run it yet. You have the prompts, the three language conditions, and the scoring all set, so running even two or three of the models you listed across the 50 prompts would give you real numbers and let you fill in those figures. That is basically a weekend of API calls, and you are closer than it probably feels. Good foundation to build on.
The project aims to produce a safety benchmark to analyze how different LLMs react to malicious requests across different types of misinformation. Specifically, the authors want to examine this in different languages and their switch variants (eg. Sheng/Swahili). Although the intention of the project is good there are little results included for proper evaluation to give a verdict for the models it tries to tackle.
Cite this work
@misc {
title={
(HckPrj) AfriSafe-CB: Evaluating LLM Safety Robustness Under African Code Switched Political and Civic Contexts
},
author={
Michelle Wanjiku Thuo, Mahmoud Mannes
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


