Evaluating Mental Health LLM Responses To Localized African English
Ewura Ama Sam, Namirah Rasul, Habeeb Abdulfatah
Large language models (LLMs) are increasingly used for informal mental health support, particularly in low-resource settings where access to professional care is limited, delayed, or costly. In such contexts, users may rely on LLMs as a first point of contact when expressing psychological distress, often using localized or non-standard English. However, most LLM safety evaluations are based on standardized English and may not reflect performance under linguistically diverse Global South communication styles.
This project investigates whether LLM responses vary in empathy, safety, relevance, and professional alignment when handling mental health prompts expressed in localized English variants. We construct a dataset of 100 base mental health prompts and generate controlled linguistic variations, including broken English, Twi–English code-mixing, and Ghanaian English usage patterns, while preserving semantic meaning. Model outputs are evaluated against clinician-authored reference responses using a structured rubric across four safety-critical dimensions. We present the experimental design and evaluation framework, and conduct systematic comparisons across linguistic conditions.
The problem statement is novel however safety mechanisms and alignment of different versions of English from context is a problem the frontier AI labs are solving / will solve, chatgpt is getting there. Mental Health responses will be met with resistence or control (similar to mythos control by the US government) and that would make this practically challenging. Though this code and dataset has wide usage in voice based models including b2b and b2c applications where I would like to see it being used from a tonality, accent and slang point of view of AI communications.
## Strengths
This is a good paper on an important and fast-growing use case. People around the world are already using LLMs as informal therapists, and that is likely to become widespread in Africa, where access to professional mental health services is severely stretched. Evaluating how a model's care quality degrades under regional language shifts is therefore a genuinely valuable thing to measure, and the team has chosen a problem that matters.
The work is well executed. The design is properly controlled, holding the semantic intent of each distress prompt constant and varying only the surface linguistic form. The use of statistical significance testing is appreciated, with paired tests reported together as a consistency check, and the results are presented clearly. The qualitative failure analysis adds real signal on top of the averages, with vivid examples such as a supportive standard-English reply turning into an outright refusal once the prompt is in broken English.
## Weaknesses
The main weakness is the limited scope. Mental health is one specific domain, so while the finding is important within it, the big-picture contribution to AI safety across the global south has a ceiling. The work is also an early signal rather than a fully validated result. It rests on a single model with manually constructed language variants.
## Recommendations for the authors
The clear recommendation is to **keep scaling this up**, because as a fully-fledged eval it could help guide the deployment of LLM-based mental health services in Africa, which would be immensely important. The most valuable steps in that direction are to **add more models** so the finding is shown to generalise, to **have native speakers validate or author the language transformations** so the conditions reflect real usage across more languge variations. Each of these turns a strong weekend signal into the kind of robust evaluation that deployers and regulators could actually rely on in Africa.
This is a great and important research direction, especially as people increasingly turn to generative LLMs for help with a wide range of concerns.
Although the research did not find any significant differences between responses to Standard English, Pidgin English and Ghanaian English prompts, except for the relevance judge, I wonder if this was because the experiment did not fully capture differences in context, environment and culture. The perturbation of the MentalChat16K dataset changes the language while preserving the underlying content. There may be sociological and cultural differences in how mental health concerns are expressed, whether they are discussed openly and what forms of support are considered appropriate. I imagine these differences would be more apparent in a multi-turn setting.
I would be interested in seeing this experiment replicated in real-world settings, using data grounded in the lived experiences of Pidgin or Ghanaian English-speaking communities.
Cite this work
@misc {
title={
(HckPrj) Evaluating Mental Health LLM Responses To Localized African English
},
author={
Ewura Ama Sam, Namirah Rasul, Habeeb Abdulfatah
},
date={
},
organization={Apart Research},
note={Research submission to the research sprint hosted by Apart.},
howpublished={https://apartresearch.com}
}


