How AI Systems Route a Query to the Right Model (Without Wasting Compute)

A modern AI system rarely uses one model for everything. It's a fleet: a cheap model for simple queries, a strong model for hard ones, maybe a retrieval-augmented specialist for factual lookups. The system routes each query to the best fit.
But routing has a hidden cost: figuring out which model fits also takes compute. Cheap estimators are noisy; accurate ones are expensive. A 2026 paper (Pandora's AI Model Routing Box) formalizes this tradeoff with surprising elegance.
The problem as a known puzzle
The researchers frame it as Pandora's Box — a classical optimization problem about costly inspection. You can open boxes to learn their value, but each opening costs something. The question: when is refining your estimate worth the cost?
Under a Gaussian signal model, they derive closed-form value-of-information expressions: for each specialist, should you spend compute to sharpen its estimated quality, or just route with what you have?
Two variants
- Centralized (Pandora's Router): a coordinator decides which estimates to refine. In benchmarks, it matched the quality of exhaustive estimation while querying the expensive estimators far less often.
- Decentralized (Pandora's Bidder): each specialist independently decides whether to invest in self-assessment. When estimates are accurate, this improves efficiency — but when they're noisy, strategic specialists can game it at others' expense.
Why this matters for agents
Model routing is becoming standard infrastructure (OpenRouter, route-based serving). This paper gives the theoretical foundation for when to spend compute deciding who should answer — a clean abstraction that could become the standard framework for cost-aware routing in production.
Someone still has to do the estimating
Pandora's Box assumes an inspection is available at a price. In a production stack that price is set by the one component nobody puts on the architecture diagram: the estimator itself. Asking a frontier model "is this query hard?" reinstates the exact cost the router existed to avoid — a second expensive call, placed in front of every cheap one.
The 2026 pattern is to make that step small on purpose. Instead of a chat call, a compact encoder reads the prompt once and returns a typed judgment — a difficulty score, a category, a calibrated yes/no — in a single forward pass. No token generation, so no streaming, no output parsing, and no second bill. The non-autoregressive decision models that climbed the Hugging Face trending board this month are the clearest version of this shape: a 421M-parameter encoder that answers typed questions over a block of state in one forward pass and returns probabilities rather than prose, at roughly 33 ms per decision. (Laya, the open-source System One decision model)
That layer is what the value-of-information math quietly leans on. Refining an estimate only pays if looking inside the box is cheap — and small encoders are what make a cheap glance possible in the first place.
For an agent orchestrating multiple models, routing isn't just "pick the biggest." It's a budget decision: how much should I spend thinking about who to ask versus just asking? That's the same planning logic that lives in every agent's architecture.
Related: How reasoning models decide how much to think · AI agent cost control · How to measure if an AI agent works · AI agent framework comparison
How reasoning models decide how much to think →AI agent cost control →how do AI agents work — return to the complete AI agent architecture guide.
Was this helpful?
Your feedback stays on this page — no tracking.