Skip to content
howdoaiagentswork.com

How AI Systems Route a Query to the Right Model (Without Wasting Compute)

How AI systems route a query to the right model

A modern AI system rarely uses one model for everything. It's a fleet: a cheap model for simple queries, a strong model for hard ones, maybe a retrieval-augmented specialist for factual lookups. The system routes each query to the best fit.

But routing has a hidden cost: figuring out which model fits also takes compute. Cheap estimators are noisy; accurate ones are expensive. A 2026 paper (Pandora's AI Model Routing Box) formalizes this tradeoff with surprising elegance.

The problem as a known puzzle

The researchers frame it as Pandora's Box — a classical optimization problem about costly inspection. You can open boxes to learn their value, but each opening costs something. The question: when is refining your estimate worth the cost?

Under a Gaussian signal model, they derive closed-form value-of-information expressions: for each specialist, should you spend compute to sharpen its estimated quality, or just route with what you have?

Two variants

Why this matters for agents

Model routing is becoming standard infrastructure (OpenRouter, route-based serving). This paper gives the theoretical foundation for when to spend compute deciding who should answer — a clean abstraction that could become the standard framework for cost-aware routing in production.

Someone still has to do the estimating

Pandora's Box assumes an inspection is available at a price. In a production stack that price is set by the one component nobody puts on the architecture diagram: the estimator itself. Asking a frontier model "is this query hard?" reinstates the exact cost the router existed to avoid — a second expensive call, placed in front of every cheap one.

The 2026 pattern is to make that step small on purpose. Instead of a chat call, a compact encoder reads the prompt once and returns a typed judgment — a difficulty score, a category, a calibrated yes/no — in a single forward pass. No token generation, so no streaming, no output parsing, and no second bill. The non-autoregressive decision models that climbed the Hugging Face trending board this month are the clearest version of this shape: a 421M-parameter encoder that answers typed questions over a block of state in one forward pass and returns probabilities rather than prose, at roughly 33 ms per decision. (Laya, the open-source System One decision model)

That layer is what the value-of-information math quietly leans on. Refining an estimate only pays if looking inside the box is cheap — and small encoders are what make a cheap glance possible in the first place.

For an agent orchestrating multiple models, routing isn't just "pick the biggest." It's a budget decision: how much should I spend thinking about who to ask versus just asking? That's the same planning logic that lives in every agent's architecture.

Related: How reasoning models decide how much to think · AI agent cost control · How to measure if an AI agent works · AI agent framework comparison

How reasoning models decide how much to think →AI agent cost control →

how do AI agents work — return to the complete AI agent architecture guide.

Was this helpful?

Your feedback stays on this page — no tracking.

Share this page