Routing LLMs using Distilled Predictors and Confidence Thresholding
Gideon Daniel Giftson T · Team Indie_Interp_Hacks
Submitted to Apart x Martian Mechanistic Router Interpretability Hackathon. Sprint projects are early-stage work by participants, not Apart Research publications.
This project explores confidence-based routing using sparsified transformer models as an intelligent alternative to monolithic AI systems. Focusing on Track 2: Intelligent Router Systems, we implemented a confidence-threshold router for a pruned DistilBERT model deployed via DeepSparse on the SST-2 sentiment classification task. We investigated how routing confidence correlates with prediction accuracy and how routing fewer, more confident samples can enable fallback to larger models while retaining accuracy. Our system enables cost-efficient, interpretable decision-making with routing justified by softmax confidence thresholds. We show that routing 70% of samples at a confidence threshold of 0.8 retains 97% of original accuracy while reducing inference costs by over 50%. These results advance the Expert Orchestration Architecture by demonstrating real-world savings and interpretable routing without compromising safety or performance.
Reviews
Great job on cost-savings and overall crisp project. I'd have looked for more robust safety considerations, like addressing confidence mis-callibration or fallback safety.
Thank you for your submission. Your submission is very readable and clear. Your paper confirms that a distilled model can perform at say 90% accuracy of its parent model, and is cheaper to run. This is a known result, reducing the novelty of this submission.
In the EO framework, the EO implementer prefers to avoid training and distilling models - leaving that to the model creators / innovators.
Distilling models for confidence based routing is a good line of work, and the ideas here are good. It would be nice to extend this, by or example - calibrating confidence of models, figuring out how to distill a LLM in different ways and so forth.
Cite this project
@misc{t2025routing,
title = {{Routing LLMs using Distilled Predictors and Confidence Thresholding}},
author = {Gideon Daniel Giftson T},
year = {2025},
month = jun,
note = {Submitted to Apart x Martian Mechanistic Router Interpretability Hackathon, an Apart Research Sprint},
howpublished = {\url{https://apartresearch.com/sprints/projects/routing-llms-using-distilled-predictors-and-confidence-thresholding-v5m6}},
url = {https://apartresearch.com/sprints/projects/routing-llms-using-distilled-predictors-and-confidence-thresholding-v5m6}
}More from Apart x Martian Mechanistic Router Interpretability Hackathon
- 1st place by peer reviewView project: Manipulating Self-Preference for Large Language Models
Manipulating Self-Preference for Large Language Models
Team Preference
Large language models (LLMs) carry great value as evaluators of synthetic data for research and production settings. However, recent research shows that language models exhibit bias towards their own responses in blind …
- 2nd place by peer reviewView project: Approximating Human Preferences Using a Multi-Judge Learned System
Approximating Human Preferences Using a Multi-Judge Learned System
AutoBox
In this work, we introduced a learned approach to aggregating multi-judge scores: using a GAM and a simple MLP as an alternative to traditional, non-learned methods like averaging. Our models outperform the naive …
- 3rd place by peer reviewView project: Judge using SAE Features
Judge using SAE Features
SAEwhat?
The key idea of this project was to explore model judgement using Sparse Autoencoder (SAE) features for mathematical reasoning tasks involving addition, multiplication, and subtraction operations. We compared this …