Skip to content
Sprint projectJun 2, 2025New Delhi, India

A generalist Router for Inspect: Reasoning router demonstration

Ishan Garg, Aman Neelappa, Gerard Boxo, Alan McBeth · Team The Interpreters

Submitted to Apart x Martian Mechanistic Router Interpretability Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.

Read the report

Report: A generalist Router for Inspect: Reasoning router demonstration

Share

With hundreds of AI models available today, choosing the right model for each query is expensive and inefficient. We develop a system that learns compact "fingerprints" of different models and automatically routes questions to the most cost-effective option. Our approach achieves 96% of premium reasoning model accuracy while cutting costs in half, making advanced AI capabilities accessible without breaking the bank. The key innovation: instead of using generic question embeddings, we let each model "see" questions through its own lens, dramatically improving routing decisions.

Reviews

Judging this Sprint?

Review this project

Your public critique appears on this page without your name. Your private critique is not published; only the Apart team reads it. If you agree below, we share your review with grantmaking.ai (opens in new tab) and the Transformative AI Fund so strong projects can be funded.

Not shown on this page.

Shown on this page, without your name.

Only the Apart team reads this, and funders if you agree below.

Share my name publicly on grantmaking.ai *
Share my private critique with funders *

  1. Thank you for your submission. I enjoyed your results - mainly a comparison of cost / benefit of two models by new criteria.

    There is no link to the code. Also the Results section starts a bit abruptly and could use more context / introduction / explanation.

  2. Strength:

    Nice problem framing saying that we don’t need extra reasoning tokens for all tasks and we can leverage a router to find the best model for each type of query; e.g: one that requires reasoning and one that does not

    Shows trade-off of cost versus accuracy for both models

    Weaknesses:

    Although we want to show that some queries don’t require reasoning, the datasets are both reasoning focussed.

    Results are a bit early stage: Ideally I would have liked to see a graph that shows the router’s performance versus every model’s performance (right now we only have performance/cost for each model per dataset)

    Could benefit from a bit of more novelty other than an application of routing to reasoning versus non-reasoning models.

    Expert Orchestration: 3

    MI : 1

    Technical Implementation and reproducibility: 1 (code is not provided)

  3. This project demonstrates an interesting interesting technical approach to model routing, and seems quite valuable from an orchestration perspective. The technical methods are clearly spelled out and seem reproducible. Though the authors didn't explicitly address safety, I could see this potentially being used to route queries to safer models for that query, and this could plausibly have impact in MI.

Cite this project

@misc{garg2025generalist,
  title = {{A generalist Router for Inspect: Reasoning router demonstration}},
  author = {Ishan Garg and Aman Neelappa and Gerard Boxo and Alan McBeth},
  year = {2025},
  month = jun,
  note = {Submitted to Apart x Martian Mechanistic Router Interpretability Hackathon, an Apart Research Sprint},
  howpublished = {\url{https://apartresearch.com/sprints/projects/a-generalist-router-for-inspect-reasoning-router-demonstration-dc46}},
  url = {https://apartresearch.com/sprints/projects/a-generalist-router-for-inspect-reasoning-router-demonstration-dc46}
}

Build something like this at the next Sprint

AI Collusion Research Sprint · Oct 23 - 25, 2026