Cross-Lingual Bias Detection in Large Language Models through Mechanistic Judge Model Evaluation
Shu Fan Sun, Fanzan Abbas, Wanjie Zhong · Team Aberdeen CLB
Submitted to Apart x Martian Mechanistic Router Interpretability Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
This paper provides a framework for detecting multilingual bias in LLMs through mechanistic interpretability of judge model behaviour. We developed an approach to evaluate language-dependent discrepancies in specialised models by leveraging semantically equivalent question-answering tasks across English, Chinese, Romanian and Vietnamese using the XQuAD dataset. Our framework utilises a judge-based evaluation system that assesses both correctness and reasoning quality on a 4-point scale, enabling the detection of biases that may not have been detected in traditional accuracy metrics alone. We demonstrate that our approach can detect these differences in model performance across languages without requiring ground truth labels, making it applicable to scenarios where traditional evaluation methods are insufficient. Our work advances Track 1 (Judge Model Development) by providing an interpretable bias detection mechanism that promotes fairness and reliability in multilingual AI systems. This framework is easily extensible to additional languages and can serve as a foundation for building fair expert orchestration systems.
Reviews
Interesting and relevant project idea to consider cross-lingual bias in LLMs. The core finding about reasoning quality varying across languages even when accuracy is similar is valuable but needs to be confirmed with stronger methodology and larger datasets.
To strengthen this work:
Add statistical analysis to support claims about performance differences; test generalization to other data sets
Validate your methodology (sanity check results from current judge prompts) and explain difference between judge 1 and judge 2 better
Group results by model to make cross-language variance easier to compare. Consider heatmaps or grouped bar charts
Motivate the research better: Consider cases where cross-lingual variance causes particular harm and think about how your findings could contribute to mitigating it (e.g. inside an expert orchestration framework)
Constructive critique:
Strength:
Bias across languages is a good motivation: a judge that detects bias accurately and robustly would be a good dimension to route to.
Varied set of languages and models to test bias/fairness.
Weakness:
I am not convinced that the rubric is detecting bias or unfair treatment.
Lack of baselines: even if the judge would be a bias detector, we would like to benchmark it against traditional approaches at bias detection.
Expert Orchestration: 3
MI: 1
Technical Imp and reproducibility: 1 (Code is a bit minimal)
Cite this project
@misc{sun2025crosslingual,
title = {{Cross-Lingual Bias Detection in Large Language Models through Mechanistic Judge Model Evaluation}},
author = {Shu Fan Sun and Fanzan Abbas and Wanjie Zhong},
year = {2025},
month = jun,
note = {Submitted to Apart x Martian Mechanistic Router Interpretability Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/crosslingual-bias-detection-in-large-language-models-through-mechanistic-judge-model-evaluation-iflu}},
url = {https://apartresearch.com/sprints/projects/crosslingual-bias-detection-in-large-language-models-through-mechanistic-judge-model-evaluation-iflu}
}More from Apart x Martian Mechanistic Router Interpretability Hackathon
- 1st place by peer reviewView project: Manipulating Self-Preference for Large Language Models
Manipulating Self-Preference for Large Language Models
Team Preference
Large language models (LLMs) carry great value as evaluators of synthetic data for research and production settings. However, recent research shows that language models exhibit bias towards their own responses in blind …
- 2nd place by peer reviewView project: Approximating Human Preferences Using a Multi-Judge Learned System
Approximating Human Preferences Using a Multi-Judge Learned System
AutoBox
In this work, we introduced a learned approach to aggregating multi-judge scores: using a GAM and a simple MLP as an alternative to traditional, non-learned methods like averaging. Our models outperform the naive …
- 3rd place by peer reviewView project: Judge using SAE Features
Judge using SAE Features
SAEwhat?
The key idea of this project was to explore model judgement using Sparse Autoencoder (SAE) features for mathematical reasoning tasks involving addition, multiplication, and subtraction operations. We compared this …