Skip to content
Sprint projectJun 2, 2025Dubai, Helsinki and California

A Multidimensional Judge Model for Safe, Consistent and Ethical AI Orchestration

Anusha Asim, Ammar Ahmed Farooqi, Sergei Smirnov, Christopher Kinoshtia · Team Sentinel!

Submitted to Apart x Martian Mechanistic Router Interpretability Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: A Multidimensional Judge Model for Safe, Consistent and Ethical AI Orchestration

Recording (opens in new tab)Code (opens in new tab)
Share

This project developed a multidimensional safety judge that transparently evaluates the outputs of large language models (LLMs) and AI routing systems across key safety dimensions: bias avoidance, factuality, manipulation resistance, and toxicity avoidance. By benchmarking both individual models and routed responses, the judge provides interpretable, auditable scores that highlight strengths, weaknesses and consistency issues. Our results show that intelligent routing, guided by these safety metrics, can deliver safer and more reliable AI outputs than monolithic models. This supports the development of trustworthy, ethical and socially responsible AI systems.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

  1. Thank you for your submission. It is well written, with some good results from comparing two base models and the Martian router.

    The matrix score diagrams would be easier to read if normalized against say the “router” values so it is easier to see better/worse. Figure descriptions need to “stand alone” and be fulsome in case the reader skim reads the paper.

  2. Comparing routed and monolithic approaches directly is a valuable direction to investigate and the team provided a good starting point with their evaluation framework!

    Some things that could make this project stronger:

    * Add statistical analysis (sample sizes, significance tests); observed differences are small enough that this is necessary

    * At least as important: Make your dataset big enough that your scores become robust

    * Validate the judge model itself through human evaluation

    * Provide dataset details when you write up your work (sample size, domains covered, generation process)

    A more ambitious and even higher-impact direction of improvement: Perform a mechanistic analysis of routing decisions to gain insights into which models were selected and why.

Cite this project

@misc{asim2025multidimensional,
  title = {{A Multidimensional Judge Model for Safe, Consistent and Ethical AI Orchestration}},
  author = {Anusha Asim and Ammar Ahmed Farooqi and Sergei Smirnov and Christopher Kinoshtia},
  year = {2025},
  month = jun,
  note = {Submitted to Apart x Martian Mechanistic Router Interpretability Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/a-multidimensional-judge-model-for-safe-consistent-and-ethical-ai-orchestration-2l0c}},
  url = {https://apartresearch.com/sprints/projects/a-multidimensional-judge-model-for-safe-consistent-and-ethical-ai-orchestration-2l0c}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026