Skip to content
Sprint projectJun 21, 2026Accra

Evaluating Mental Health LLM Responses To Localized African English

Ewura Ama Sam, Namirah Rasul, Habeeb Abdulfatah · Team HENA

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Evaluating Mental Health LLM Responses To Localized African English

Code (opens in new tab)
Share

Large language models (LLMs) are increasingly used for informal mental health support, particularly in low-resource settings where access to professional care is limited, delayed, or costly. In such contexts, users may rely on LLMs as a first point of contact when expressing psychological distress, often using localized or non-standard English. However, most LLM safety evaluations are based on standardized English and may not reflect performance under linguistically diverse Global South communication styles. This project investigates whether LLM responses vary in empathy, safety, relevance, and professional alignment when handling mental health prompts expressed in localized English variants. We construct a dataset of 100 base mental health prompts and generate controlled linguistic variations, including broken English, Twi–English code-mixing, and Ghanaian English usage patterns, while preserving semantic meaning. Model outputs are evaluated against clinician-authored reference responses using a structured rubric across four safety-critical dimensions. We present the experimental design and evaluation framework, and conduct systematic comparisons across linguistic conditions.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The problem statement is novel however safety mechanisms and alignment of different versions of English from context is a problem the frontier AI labs are solving / will solve, chatgpt is getting there. Mental Health responses will be met with resistence or control (similar to mythos control by the US government) and that would make this practically challenging. Though this code and dataset has wide usage in voice based models including b2b and b2c applications where I would like to see it being used from a tonality, accent and slang point of view of AI communications.

  2. This is a great and important research direction, especially as people increasingly turn to generative LLMs for help with a wide range of concerns.

    Although the research did not find any significant differences between responses to Standard English, Pidgin English and Ghanaian English prompts, except for the relevance judge, I wonder if this was because the experiment did not fully capture differences in context, environment and culture. The perturbation of the MentalChat16K dataset changes the language while preserving the underlying content. There may be sociological and cultural differences in how mental health concerns are expressed, whether they are discussed openly and what forms of support are considered appropriate. I imagine these differences would be more apparent in a multi-turn setting.

    I would be interested in seeing this experiment replicated in real-world settings, using data grounded in the lived experiences of Pidgin or Ghanaian English-speaking communities.

    Read full reviewShow less
  3. ## Strengths

    This is a good paper on an important and fast-growing use case. People around the world are already using LLMs as informal therapists, and that is likely to become widespread in Africa, where access to professional mental health services is severely stretched. Evaluating how a model's care quality degrades under regional language shifts is therefore a genuinely valuable thing to measure, and the team has chosen a problem that matters.

    The work is well executed. The design is properly controlled, holding the semantic intent of each distress prompt constant and varying only the surface linguistic form. The use of statistical significance testing is appreciated, with paired tests reported together as a consistency check, and the results are presented clearly. The qualitative failure analysis adds real signal on top of the averages, with vivid examples such as a supportive standard-English reply turning into an outright refusal once the prompt is in broken English.

    ## Weaknesses

    The main weakness is the limited scope. Mental health is one specific domain, so while the finding is important within it, the big-picture contribution to AI safety across the global south has a ceiling. The work is also an early signal rather than a fully validated result. It rests on a single model with manually constructed language variants.

    ## Recommendations for the authors

    The clear recommendation is to **keep scaling this up**, because as a fully-fledged eval it could help guide the deployment of LLM-based mental health services in Africa, which would be immensely important. The most valuable steps in that direction are to **add more models** so the finding is shown to generalise, to **have native speakers validate or author the language transformations** so the conditions reflect real usage across more languge variations. Each of these turns a strong weekend signal into the kind of robust evaluation that deployers and regulators could actually rely on in Africa.

    Read full reviewShow less

Cite this project

@misc{sam2026evaluating,
  title = {{Evaluating Mental Health LLM Responses To Localized African English}},
  author = {Ewura Ama Sam and Namirah Rasul and Habeeb Abdulfatah},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/evaluating-mental-health-llm-responses-to-localized-african-english-pzag}},
  url = {https://apartresearch.com/sprints/projects/evaluating-mental-health-llm-responses-to-localized-african-english-pzag}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026