Skip to content
Sprint projectJun 2, 2025Aberdeen

Cross-Lingual Bias Detection in Large Language Models through Mechanistic Judge Model Evaluation

Shu Fan Sun, Fanzan Abbas, Wanjie Zhong · Team Aberdeen CLB

Submitted to Apart x Martian Mechanistic Router Interpretability Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: Cross-Lingual Bias Detection in Large Language Models through Mechanistic Judge Model Evaluation

Code (opens in new tab)
Share

This paper provides a framework for detecting multilingual bias in LLMs through mechanistic interpretability of judge model behaviour. We developed an approach to evaluate language-dependent discrepancies in specialised models by leveraging semantically equivalent question-answering tasks across English, Chinese, Romanian and Vietnamese using the XQuAD dataset. Our framework utilises a judge-based evaluation system that assesses both correctness and reasoning quality on a 4-point scale, enabling the detection of biases that may not have been detected in traditional accuracy metrics alone. We demonstrate that our approach can detect these differences in model performance across languages without requiring ground truth labels, making it applicable to scenarios where traditional evaluation methods are insufficient. Our work advances Track 1 (Judge Model Development) by providing an interpretable bias detection mechanism that promotes fairness and reliability in multilingual AI systems. This framework is easily extensible to additional languages and can serve as a foundation for building fair expert orchestration systems.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

  1. Interesting and relevant project idea to consider cross-lingual bias in LLMs. The core finding about reasoning quality varying across languages even when accuracy is similar is valuable but needs to be confirmed with stronger methodology and larger datasets.

    To strengthen this work:

    Add statistical analysis to support claims about performance differences; test generalization to other data sets

    Validate your methodology (sanity check results from current judge prompts) and explain difference between judge 1 and judge 2 better

    Group results by model to make cross-language variance easier to compare. Consider heatmaps or grouped bar charts

    Motivate the research better: Consider cases where cross-lingual variance causes particular harm and think about how your findings could contribute to mitigating it (e.g. inside an expert orchestration framework)

  2. Constructive critique:

    Strength:

    Bias across languages is a good motivation: a judge that detects bias accurately and robustly would be a good dimension to route to.

    Varied set of languages and models to test bias/fairness.

    Weakness:

    I am not convinced that the rubric is detecting bias or unfair treatment.

    Lack of baselines: even if the judge would be a bias detector, we would like to benchmark it against traditional approaches at bias detection.

    Expert Orchestration: 3

    MI: 1

    Technical Imp and reproducibility: 1 (Code is a bit minimal)

Cite this project

@misc{sun2025crosslingual,
  title = {{Cross-Lingual Bias Detection in Large Language Models through Mechanistic Judge Model Evaluation}},
  author = {Shu Fan Sun and Fanzan Abbas and Wanjie Zhong},
  year = {2025},
  month = jun,
  note = {Submitted to Apart x Martian Mechanistic Router Interpretability Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/crosslingual-bias-detection-in-large-language-models-through-mechanistic-judge-model-evaluation-iflu}},
  url = {https://apartresearch.com/sprints/projects/crosslingual-bias-detection-in-large-language-models-through-mechanistic-judge-model-evaluation-iflu}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026