Skip to content
Sprint projectJun 21, 2026Cape Town

Ufakazi

Kevin Brand, Racquel Dennison · Team Ufakazi

Submitted to Global South AI Safety Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Ufakazi is an evaluation harness built to determine whether models are unfairly biased towards trusting testimonies presented in high-resource languages (specifically English). By controlling for confounders, we isolate the language bias of various current generation LLMs and show that they often unfairly discriminate against marginalised and underrepresented communities.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

How much would this matter for AI safety if it worked? How innovative is it? For scores of 4-5: is this actually new to the field, or replicating recent work?

Scoring guide
  1. 1Negligible. No clear problem addressed, or no meaningful novelty.
  2. 2Limited. Addresses a real problem but with a generic or well-trodden approach. Incremental at best.
  3. 3Moderate. Clear problem with a reasonable approach; some novelty in framing or method beyond routine application of existing tools.
  4. 4Significant. Important problem with an original approach, or identifies a neglected problem area. A valuable contribution others could build on.
  5. 5Exceptional. Tackles a critical AI safety problem with a genuinely novel approach, or opens a new research direction. Clear theory of change. You'd be excited to share this with researchers in the area.

How sound are methodology, implementation, and findings?

Scoring guide
  1. 1Seriously flawed. Methodology broken, results uninterpretable, or implementation doesn't work.
  2. 2Weak. Approach has significant gaps: missing validation, flawed experimental design, or incomplete implementation.
  3. 3Competent. Technically solid given the short duration. Methodology makes sense, results are interpretable, limitations acknowledged, work builds toward clear conclusions.
  4. 4Strong. Thorough methodology with convincing validation. Results clearly support conclusions. Immediately useful for future work.
  5. 5Exceptional. Ambitious scope executed rigorously. Surprising findings, novel methods, or unusually robust validation.

How clearly are work, findings, and impact potential communicated?

Scoring guide
  1. 1Incomprehensible. Cannot determine what the project is actually claiming or doing.
  2. 2Hard to follow. Key information buried, missing, or diluted by excessive length. Significant effort to extract main points.
  3. 3Clear enough. Can understand the problem, approach, and results without undue effort. Core content clearly present: problem, method, findings, limitations.
  4. 4Well presented. Easy to follow, well-structured, appropriate level of detail. Target audience would get it quickly.
  5. 5Exceptionally clear. A pleasure to read. Complex ideas made accessible. Could serve as a model for how to present this type of work.

  1. ## Strengths

    This is an important finding, demonstrated well. LLMs are already being deployed in legal settings, so a credibility bias that tracks the language a testimony is written in is a concrete, high-stakes risk rather than a hypothetical one. The work shows that 7 of 9 models systematically prefer English testimony over evidentially-balanced testimony in Afrikaans, isiZulu, or isiXhosa, and that the effect deepens as the language gets lower-resourced.

    The strongest contribution is the effort to isolate language discrimination as the sole cause. The design holds the content and the answer position fixed and varies only the language, with same-language controls to confirm the setup is neutral before trusting the cross-language numbers. The team then closes off the obvious alternative explanations: a human-versus-machine translation provenance check, a verbosity and length check, and a Gemma scale ladder. The rationale analysis is a standout, because the models often name the language of the testimony as their reason, which turns a statistical signal into interpretable evidence.

    ## Weaknesses

    The finding is important, but its impact has a ceiling. It demonstrates a risk in one deployment context rather than mitigating it or opening a broad new direction, which keeps it at significant rather than exceptional. The scope is also bounded by the data. The scenario set is small, so some confidence intervals are wide and the results are indicative rather than conclusive on the most uncertain cells. The two models closest to the frontier did not show the bias, and the paper does not yet establish why, which leaves the most decision-relevant question open.

    ## Recommendations for the authors

    This is strong, well-executed work, and a few extensions would turn it into a full paper. First, **expand the scenario set so the confidence intervals can be computed reliably**, which would move the borderline results from indicative to conclusive. Second, **evaluate the frontier models directly**. The signal that the strongest models may not share the bias is the most important open question, and a full version needs to establish whether that holds. If it does, you can give legal practitioners a concrete, evidence-based recommendation: that weaker models cannot be trusted for high-stakes testimonial work, and which classes of model are safer to deploy.

    Read full reviewShow less
  2. The findings in this project carry clear implications and the authors can be commended for the problem they have identified. I did however find the explanation quite lengthy and perhaps a bit too technical / over-explained - the introduction was clear and then it became difficult to follow until 'The bias is in the reasoning, not only the choice...' (page 9). It certainly is something that courts should be aware of in relying on LLMs to assist with witness testimonies (or any documentation referred to the court in a language other than English). Well thought out.

  3. This was a well-executed project with a thoughtful methodology, especially given the time constraints. I appreciated the use of clear controls and validity checks which made the results easier to interpret. I also appreciated the thoroughness in ruling out other confounding factors.

    The methodology felt especially strong in how it separated different possible sources of bias.

    Overall, this was a rigorous and well-scoped.

  4. The project is well-executed and addresses an important problem. However, there are two notable limitations:

    - The evaluation is based on general-purpose language models rather than models specifically designed or fine-tuned for legal tasks, limiting the strength of the conclusions about legal reasoning.

    - The paper does not attempt to characterize the likely distribution of the models' training data across languages or regions. While frontier model developers do not disclose their training data composition, it is reasonable to expect these datasets to be heavily skewed toward high-resource languages. This makes it difficult to determine whether the observed biases are inherent to the models or primarily a consequence of underlying data imbalances.

Cite this project

@misc{brand2026ufakazi,
  title = {{Ufakazi}},
  author = {Kevin Brand and Racquel Dennison},
  year = {2026},
  month = jun,
  note = {Submitted to Global South AI Safety Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/ufakazi-7yi3}},
  url = {https://apartresearch.com/sprints/projects/ufakazi-7yi3}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026