Skip to content
Sprint projectJun 21, 2026St. Paul, Minnesota, USA

AfriSafeBench: Evaluating LLM Recognition of AI Safety and Governance Risks in African Healthcare AI Deployments

Kimberly Atta-Peters · Team AfriSafeBench

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: AfriSafeBench: Evaluating LLM Recognition of AI Safety and Governance Risks in African Healthcare AI Deployments

Code (opens in new tab)
Share

AfriSafeBench is a benchmark and tool for testing whether LLMs can identify AI safety and governance risks in African healthcare AI deployment scenarios.

I created 25 scenarios across seven African countries and evaluated three models: llama-3.1-8b-instant, llama-3.3-70b-versatile, and openai/gpt-oss-20b. The models were scored on whether they detected expected risks such as bias, human oversight, local validation, privacy, vendor dependency, resource constraints, and unsafe medical advice.

The results showed that models often detected familiar risks like bias and privacy, but struggled with structural deployment risks such as resource-constrained deployment, vendor dependency, monitoring, and local validation. AfriSafeBench also includes a framework-guided report mode that uses WHO, NIST, UNESCO, OECD, and African Union guidance to generate governance recommendations.

The project shows where LLMs can support AI safety review, where they miss important risks, and why human oversight remains necessary in healthcare AI governance.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. The framing is the strongest part: separating recognition of structural deployment risks (resource constraints, vendor dependency, local validation, monitoring) from familiar AI-ethics knowledge like bias and privacy is a genuinely useful distinction, and the African healthcare focus is underserved.

    My main concern is on execution. 25 scenarios x 3 models = 75 evaluations is too small to support the model-comparison claims; the report is honest about this ("descriptive rather than statistically significant"), and the 100-point spread on Scenario 002 looks like a single-scenario artifact at this n rather than a model property. The bigger issue is that 40 of 75 evaluations (53.3%) needed human rescoring because automated scoring missed semantic matches, but those human judgments come from a single rater with no inter-rater agreement / Cohen's kappa and no domain-expert validation of the researcher-defined expected categories. That makes the "ground truth" one person's call. There's also no external baseline to anchor the 72% coverage figure. And the multilingual angle is in tension with the Limitations section's own statement that all 25 scenarios are English-language.

    Presentation is clear — the category tables, Figure 1, and the workflow diagram communicate well, and the Limitations/Dual-Use sections are candidly written.

    Most valuable next step: have 2-3 African clinical or policy reviewers independently re-score a scenario subset and report agreement. That single move converts this from one researcher's labels into a defensible benchmark.

    Read full reviewShow less
  2. Strong and carefully run for a hackathon. The best part is that it tests whether models catch the "hidden" risks — funding running out, no local testing, vendor lock-in — and shows they miss these while spotting obvious bias and privacy issues. The scoring was careful and honest, openly flagging that 53% of cases needed a human check. To improve: have African health and policy experts confirm the "expected risks," since you defined them yourself, and add a second human rater. The report mode is a nice extra, but barely tested.

  3. This project identifies a highly neglected problem area: structural and governance risks in Global South healthcare deployments, as opposed to standard AI ethics. The integration of a scenario-based benchmark with a framework-guided governance report mode using WHO, NIST, and OECD documents is genuinely novel and impactful. The insight that models perform well on familiar ethics concepts but fail to recognize structural risks like resource constraints and vendor dependency is profound. To improve the execution score, expanding the dataset beyond 25 English-language scenarios to include multilingual or local-language contexts would strengthen the benchmark's robustness.

Cite this project

@misc{attapeters2026afrisafebench,
  title = {{AfriSafeBench: Evaluating LLM Recognition of AI Safety and Governance Risks in African Healthcare AI Deployments}},
  author = {Kimberly Atta-Peters},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/afrisafebench-evaluating-llm-recognition-of-ai-safety-and-governance-risks-in-african-healthcare-ai-deployments-k2iz}},
  url = {https://apartresearch.com/sprints/projects/afrisafebench-evaluating-llm-recognition-of-ai-safety-and-governance-risks-in-african-healthcare-ai-deployments-k2iz}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026