Skip to content
Sprint projectJun 22, 2026Lagos, Nigeria

Benchmarking Open-Weight vs. Frontier LLMs on African Health and Financial-Inclusion Reasoning, With and Without Graph RAG

Chibuokem Faithful Chukwunwogor · Team Faithful

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Benchmarking Open-Weight vs. Frontier LLMs on African Health and Financial-Inclusion Reasoning, With and Without Graph RAG

Code (opens in new tab)
Share

This study tested whether lightweight open-weight LLMs (Qwen3.6-27B, Gemma-4-31B-it), with Graph RAG grounding, could rival frontier models (GPT-5, Gemini Pro) on African health and financial reasoning—reducing compute dependency and data-colonial reliance on foreign infrastructure. GPT-5 led both domains; Qwen3.6-27B without RAG was the strongest open-weight performer, approaching GPT-5 on health (0.81 vs 0.88). Critically, RAG *degraded* accuracy in three of four open-weight conditions, undermining the core hypothesis. Gemini Pro's near-zero scores likely reflect a pipeline artifact. Small samples, API-based (not local) inference, and scarce African financial-LLM benchmarks limit conclusions and motivate further research.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. ## Strengths

    This is a useful capability study in two domains that matter a great deal for African AI services: healthcare and financial inclusion. Improving the quality and reliability of AI in these areas is genuinely valuable, and grounding the work in real African data (AfriMed-QA and FinScope) keeps it relevant to the region rather than generic.

    The team built a real end-to-end pipeline rather than just describing one, and they are intellectually honest about the outcome. The headline result is negative: Graph RAG degraded performance in most of the open-weight conditions rather than improving it. Reporting that clearly, instead of quietly dropping it, is the right thing to do and makes the finding more trustworthy. The sovereignty motivation also connects the work to real regional concerns about compute scarcity and dependency on foreign infrastructure.

    ## Weaknesses

    The main limitation is that the AI safety relevance is modest at the big-picture scale. This reads more as a capability and service-quality study than as a contribution to AI safety specifically, which lowers its impact within this track. The intervention is also a reasonable but well-established choice. Graph RAG is tried and tested, so the technical novelty is limited, and the most interesting result here is the negative one rather than a new method. Taken together, the work is valuable as applied engineering but more limited as a research contribution.

    ## Recommendations for the authors

    The most direct way to realise this work's value is to **share the findings and ideas with companies and teams actively deploying health and finance AI in African markets**. The negative Graph RAG result is a practically useful, cost-saving signal for practitioners, and that audience is where it will have the most immediate impact.

    Read full reviewShow less
  2. The project did well in clearly presenting the domains being studied, healthcare and financial services, and explaining why they matter.

    The paper also does a solid job interpreting the patterns found in the experiment. The discussion of surprising results, such as Gemini Pro’s near-zero scores, was thoughtful, and the authors did a good job forming hypotheses around unexpected model behavior rather than just reporting the numbers.

    One area that felt less fully resolved was the project’s central sovereignty question: “Can African research institutions retain meaningful control over the AI systems mediating health and financial decisions affecting their populations?” The paper acknowledges that sovereignty could not be fully tested because the experiment substituted API-hosted Qwen/Gemma models for local GGUF inference. Because of this, the study seems better positioned as an initial benchmark of model performance and Graph RAG effects, rather than a full answer to the sovereignty question.

    Read full reviewShow less
  3. Strengths: Explores a technically relevant topic with strong practical implications for enterprise AI systems. The comparison between approaches is well motivated and clearly presented. Areas for Improvement: Strengthen experimental rigor by expanding benchmark datasets, documenting evaluation methodology, and including statistical comparisons across models. More discussion of trade-offs such as cost, latency, accuracy, and operational complexity would improve the overall analysis.

Cite this project

@misc{chukwunwogor2026benchmarking,
  title = {{Benchmarking Open-Weight vs. Frontier LLMs on African Health and Financial-Inclusion Reasoning, With and Without Graph RAG}},
  author = {Chibuokem Faithful Chukwunwogor},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/benchmarking-openweight-vs-frontier-llms-on-african-health-and-financialinclusion-reasoning-with-and-without-graph-rag-xel1}},
  url = {https://apartresearch.com/sprints/projects/benchmarking-openweight-vs-frontier-llms-on-african-health-and-financialinclusion-reasoning-with-and-without-graph-rag-xel1}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026