SomaliCrowS: A Benchmark for Evaluating Gender Bias in Large Language Models Using the Somali Language
Abdullahi Hassan · Team EA Somalia AI Safety Lab
Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
Large language models are now used in education, healthcare, and public services across Somali-speaking communities. But most bias-testing benchmarks are built for English, so we don't know if these models treat Somali text fairly. We introduce SomaliCrowS, the first benchmark for measuring gender bias in language models using the Somali language. Following the CrowS-Pairs method, we built 220 sentence pairs across seven categories, Politics, Business, Leadership, STEM, Family, Occupation, and Education, where the only difference between sentences is the subject's gender. We tested XLM-RoBERTa on these pairs by comparing how likely the model thought each version was. The model favored the male version in 87.3% of all pairs, and every category showed a statistically significant departure from an unbiased 50/50 split (binomial tests, p < 0.05). Bias was most pronounced in Leadership (97.5% male-preferred) and Politics (largest mean log-probability gap, −7.096), and weakest in Education (67.5%, −0.164). These results show clear gender bias in a widely used multilingual model when it processes Somali. SomaliCrowS gives researchers a reusable tool to catch this kind of bias before deploying AI in Somali-speaking communities.
Reviews
On Impact Potential and Innovation, it was quite strong. To push this toward an "Exceptional" (5), future iterations of the work could introduce a genuinely novel evaluation method tailored specifically to the linguistic or cultural nuances of Somali, rather than relying exclusively on a Western-developed framework like CrowS-Pairs.
You built the first way to measure gender bias in AI models for the Somali language, which matters because these models already get used in Somali schools and services with no bias check today. The execution is clean, and you published the data and notebook so others can reuse it. The leadership and politics results are believable. Two things to be aware of. The method is a direct adaptation of an existing English benchmark, so the contribution is really the language, not the technique. And you measure the model's internal preference rather than what a real user would see in a reply, so I'd run the same tests on a model people actually chat with. Gender is a good start. Clan and ethnicity can carry Somali bias and would make this far more useful.
SomaliCrowS fills a real gap and the CrowS-Pairs adaptation is methodologically appropriate and clearly executed. The most important next step is formal community validation. The benchmark is built on one researcher's judgment about culturally salient stereotypes
Cite this project
@misc{hassan2026somalicrows,
title = {{SomaliCrowS: A Benchmark for Evaluating Gender Bias in Large Language Models Using the Somali Language}},
author = {Abdullahi Hassan},
year = {2026},
month = jun,
note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/somalicrows-a-benchmark-for-evaluating-gender-bias-in-large-language-models-using-the-somali-language-xw3c}},
url = {https://apartresearch.com/sprints/projects/somalicrows-a-benchmark-for-evaluating-gender-bias-in-large-language-models-using-the-somali-language-xw3c}
}More from Global South AI Safety Hackathon
- View project: Pragmatic Sophistry in Vietnamese Multi-Agent Oversight
Pragmatic Sophistry in Vietnamese Multi-Agent Oversight
AI Safety Enthusiasts
AI safety monitors are usually evaluated on the assumption that risky behavior is lexically visible in the text being watched. We test this assumption in a multilingual, multi-agent setting: Vietnamese-language workflow …
- View project: JusticIA: A Counterfactual Benchmark for Auditing Contextual Biases in Language Models for Transitional Justice
JusticIA: A Counterfactual Benchmark for Auditing Contextual Biases in Language Models for Transitional Justice
JusticeMiners
JusticIA is a counterfactual benchmark for auditing contextual bias in LLMs applied to Colombian transitional justice. It tests whether six LLMs change their sanction recommendations when only one contextual attribute …
- View project: Coldron
Coldron
ColDron
En Colombia, los grupos armados ilegales ya atacan con drones comerciales modificados y ya han herido y matado a civiles. Una pregunta decide cómo gobernar esta amenaza: ¿quién elige el blanco y aprieta el gatillo? Hoy, …