A generalist Router for Inspect: Reasoning router demonstration
Ishan Garg, Aman Neelappa, Gerard Boxo, Alan McBeth · Team The Interpreters
Submitted to Apart x Martian Mechanistic Router Interpretability Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
With hundreds of AI models available today, choosing the right model for each query is expensive and inefficient. We develop a system that learns compact "fingerprints" of different models and automatically routes questions to the most cost-effective option. Our approach achieves 96% of premium reasoning model accuracy while cutting costs in half, making advanced AI capabilities accessible without breaking the bank. The key innovation: instead of using generic question embeddings, we let each model "see" questions through its own lens, dramatically improving routing decisions.
Reviews
Thank you for your submission. I enjoyed your results - mainly a comparison of cost / benefit of two models by new criteria.
There is no link to the code. Also the Results section starts a bit abruptly and could use more context / introduction / explanation.
Strength:
Nice problem framing saying that we don’t need extra reasoning tokens for all tasks and we can leverage a router to find the best model for each type of query; e.g: one that requires reasoning and one that does not
Shows trade-off of cost versus accuracy for both models
Weaknesses:
Although we want to show that some queries don’t require reasoning, the datasets are both reasoning focussed.
Results are a bit early stage: Ideally I would have liked to see a graph that shows the router’s performance versus every model’s performance (right now we only have performance/cost for each model per dataset)
Could benefit from a bit of more novelty other than an application of routing to reasoning versus non-reasoning models.
Expert Orchestration: 3
MI : 1
Technical Implementation and reproducibility: 1 (code is not provided)
This project demonstrates an interesting interesting technical approach to model routing, and seems quite valuable from an orchestration perspective. The technical methods are clearly spelled out and seem reproducible. Though the authors didn't explicitly address safety, I could see this potentially being used to route queries to safer models for that query, and this could plausibly have impact in MI.
Cite this project
@misc{garg2025generalist,
title = {{A generalist Router for Inspect: Reasoning router demonstration}},
author = {Ishan Garg and Aman Neelappa and Gerard Boxo and Alan McBeth},
year = {2025},
month = jun,
note = {Submitted to Apart x Martian Mechanistic Router Interpretability Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/a-generalist-router-for-inspect-reasoning-router-demonstration-dc46}},
url = {https://apartresearch.com/sprints/projects/a-generalist-router-for-inspect-reasoning-router-demonstration-dc46}
}More from Apart x Martian Mechanistic Router Interpretability Hackathon
- 1st place by peer reviewView project: Manipulating Self-Preference for Large Language Models
Manipulating Self-Preference for Large Language Models
Team Preference
Large language models (LLMs) carry great value as evaluators of synthetic data for research and production settings. However, recent research shows that language models exhibit bias towards their own responses in blind …
- 2nd place by peer reviewView project: Approximating Human Preferences Using a Multi-Judge Learned System
Approximating Human Preferences Using a Multi-Judge Learned System
AutoBox
In this work, we introduced a learned approach to aggregating multi-judge scores: using a GAM and a simple MLP as an alternative to traditional, non-learned methods like averaging. Our models outperform the naive …
- 3rd place by peer reviewView project: Judge using SAE Features
Judge using SAE Features
SAEwhat?
The key idea of this project was to explore model judgement using Sparse Autoencoder (SAE) features for mathematical reasoning tasks involving addition, multiplication, and subtraction operations. We compared this …