Skip to content
Sprint projectJun 20, 2026London

ParentMe SafeAI: An African Child and Family AI Safety Evaluation Framework

Rita · Team Rita Zadi

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: ParentMe SafeAI: An African Child and Family AI Safety Evaluation Framework

Share

ParentMe SafeAI: An African Child and Family AI Safety Evaluation Framework

ParentMe SafeAI is a project that aims to improve AI safety for children, parents, and families across Africa by creating an evaluation framework that tests how AI systems respond to real-world family and child welfare challenges.

Current AI safety benchmarks are largely developed using Western datasets and often fail to capture risks that are relevant to African communities. ParentMe SafeAI addresses this gap by developing a set of safety tests and evaluation criteria focused on child safeguarding, parenting advice, health misinformation, educational guidance, financial scams targeting families, and bias against vulnerable groups.

The project will create a child and family harm taxonomy, a multilingual dataset of evaluation prompts based on African contexts, and a scoring framework to assess whether AI systems provide safe, accurate, culturally appropriate, and responsible responses.

By centring African realities and family wellbeing, ParentMe SafeAI seeks to help developers, policymakers, and organisations build AI systems that are safer, more inclusive, and more beneficial for children and families across the continent.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. This project identifies a critical and frequently overlooked gap in AI safety: the cultural, linguistic, and socio-economic realities of African communities. The focus on localized harms—such as grant scams, malaria misinformation, and the nuances of multi-generational households—is highly insightful and urgently needed. The taxonomy is thoroughly planned out.

    To take this project to the next level, consider the following pointers:

    - Scale the Dataset: An evaluation dataset of 50-100 prompts is an excellent starting point for a hackathon proof-of-concept. However, for a robust benchmark, this will need to be significantly larger. Consider proposing a strategy for crowdsourcing or synthetically generating localized prompts.

    - Define the Scoring Mechanism: The proposal outlines the "Safety Scoring Framework" as an expected output, but lacks details on how the scoring will actually work. Specifying whether the framework will use human annotators, LLM-as-a-judge, or rule-based metrics would greatly strengthen the methodological soundness.

    - Demonstrate a Proof of Concept: Currently, the document reads primarily as a strong proposal rather than an executed project. Including a small pilot test—such as running 5 of these localized prompts through a popular LLM and showing where it fails—would provide concrete "findings" and significantly boost the execution quality.

    - Linguistic Nuance: While incorporating languages like Swahili and Hausa is slated for future versions, you could suggest that even their initial English prompts should account for regional dialects, slang, or vernacular (e.g., West African Pidgin) to test real-world user interactions.

    Read full reviewShow less
  2. The main limitation is that it is still mostly at the idea stage. To make it stronger, I would suggest building a small first version of the benchmark, with example prompts, scoring guidelines, and a simple test on one or two AI systems.

  3. Strong problem, and you framed it well. Western built safety tests really do miss the dangers that matter most to African families, and your five categories map that gap clearly and thoughtfully. The honest gap is that this is a plan rather than a piece of work. There is no dataset, no prompts run, and no results, so there is not yet anything to evaluate beyond the framing. The thing that would have moved this furthest is a small slice built end to end, even ten or twenty prompts run against one model with your scoring applied. I would build that minimum version next and show what it actually catches, since the moment you have real examples of a model giving unsafe advice in these settings, the case makes itself.

Cite this project

@misc{rita2026parentme,
  title = {{ParentMe SafeAI: An African Child and Family AI Safety Evaluation Framework}},
  author = {Rita},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/parentme-safeai-an-african-child-and-family-ai-safety-evaluation-framework-ncl9}},
  url = {https://apartresearch.com/sprints/projects/parentme-safeai-an-african-child-and-family-ai-safety-evaluation-framework-ncl9}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026