A Multidimensional Judge Model for Safe, Consistent and Ethical AI Orchestration
Anusha Asim, Ammar Ahmed Farooqi, Sergei Smirnov, Christopher Kinoshtia · Team Sentinel!
Submitted to Apart x Martian Mechanistic Router Interpretability Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
This project developed a multidimensional safety judge that transparently evaluates the outputs of large language models (LLMs) and AI routing systems across key safety dimensions: bias avoidance, factuality, manipulation resistance, and toxicity avoidance. By benchmarking both individual models and routed responses, the judge provides interpretable, auditable scores that highlight strengths, weaknesses and consistency issues. Our results show that intelligent routing, guided by these safety metrics, can deliver safer and more reliable AI outputs than monolithic models. This supports the development of trustworthy, ethical and socially responsible AI systems.
Reviews
Thank you for your submission. It is well written, with some good results from comparing two base models and the Martian router.
The matrix score diagrams would be easier to read if normalized against say the “router” values so it is easier to see better/worse. Figure descriptions need to “stand alone” and be fulsome in case the reader skim reads the paper.
Comparing routed and monolithic approaches directly is a valuable direction to investigate and the team provided a good starting point with their evaluation framework!
Some things that could make this project stronger:
* Add statistical analysis (sample sizes, significance tests); observed differences are small enough that this is necessary
* At least as important: Make your dataset big enough that your scores become robust
* Validate the judge model itself through human evaluation
* Provide dataset details when you write up your work (sample size, domains covered, generation process)
A more ambitious and even higher-impact direction of improvement: Perform a mechanistic analysis of routing decisions to gain insights into which models were selected and why.
Cite this project
@misc{asim2025multidimensional,
title = {{A Multidimensional Judge Model for Safe, Consistent and Ethical AI Orchestration}},
author = {Anusha Asim and Ammar Ahmed Farooqi and Sergei Smirnov and Christopher Kinoshtia},
year = {2025},
month = jun,
note = {Submitted to Apart x Martian Mechanistic Router Interpretability Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/a-multidimensional-judge-model-for-safe-consistent-and-ethical-ai-orchestration-2l0c}},
url = {https://apartresearch.com/sprints/projects/a-multidimensional-judge-model-for-safe-consistent-and-ethical-ai-orchestration-2l0c}
}More from Apart x Martian Mechanistic Router Interpretability Hackathon
- 1st place by peer reviewView project: Manipulating Self-Preference for Large Language Models
Manipulating Self-Preference for Large Language Models
Team Preference
Large language models (LLMs) carry great value as evaluators of synthetic data for research and production settings. However, recent research shows that language models exhibit bias towards their own responses in blind …
- 2nd place by peer reviewView project: Approximating Human Preferences Using a Multi-Judge Learned System
Approximating Human Preferences Using a Multi-Judge Learned System
AutoBox
In this work, we introduced a learned approach to aggregating multi-judge scores: using a GAM and a simple MLP as an alternative to traditional, non-learned methods like averaging. Our models outperform the naive …
- 3rd place by peer reviewView project: Judge using SAE Features
Judge using SAE Features
SAEwhat?
The key idea of this project was to explore model judgement using Sparse Autoencoder (SAE) features for mathematical reasoning tasks involving addition, multiplication, and subtraction operations. We compared this …