- Within one vendor, the cheapest and most expensive current models differ 40-fold (Claude Haiku 5.5 vs Opus 5.5) to 100-fold (GPT-6 Luna vs GPT-6 Astra) in list price.
- Routing patterns include a classifier that picks a model up front, a cascade that escalates when a cheap answer fails, and a planner or advisor that brings in a strong model only where it matters.
- Vendor features are narrower than the name suggests. Anthropic's advisor tool is in beta, its 'fallbacks' only handle safety refusals, and OpenAI's GPT-6 guide describes no automatic model selection.
Contents
Most production AI traffic is easy: tagging a ticket, pulling a date out of an email, summarizing a page. A small amount is hard. Model routing means sending each request to the cheapest model that can handle it. In October 2026 it's one of the biggest cost levers left, because the gap between small models and frontier models within a single vendor is huge. On Anthropic's official pricing page, Claude Haiku 5.5 costs $0.10 per million input tokens and $0.50 per million output tokens (for prompts up to 100K tokens), while Claude Opus 5.5 costs $4 and $20. That's a 40x gap. On OpenAI's, GPT-6 Luna costs $0.10/$0.50 and GPT-6 Astra costs $10/$50, a 100x gap.
This explainer covers how routing works, the main patterns, what Anthropic, OpenAI and OpenRouter actually ship (checked against their docs on October 8, 2026), and what the savings look like on paper.
The terminology
- Router: a component that decides, before the work starts, which model handles a request. It can be rules, a small classifier, or another LLM.
- Cascade: try the cheap model first, check the result, and escalate to a stronger model if it isn't good enough.
- Planner/executor: a strong model plans and a cheaper model carries out the steps. Some setups use the reverse: a cheap executor that consults a strong advisor.
- Subagent: a separate agent instance, often on a smaller model, that handles a bounded subtask (search, summarize, explore) and returns a compact result to the main agent.
- Fallback: retry on another model when the first one fails. "Fails" can mean an error, a rate limit or a refusal, depending on the system.
Pattern 1: a classifier router
A router looks at the request and picks a tier. It can be as simple as a rule ("all requests from the tagging endpoint go to Haiku") or a learned model that predicts difficulty.
The best-known research here is RouteLLM (Ong et al., arXiv 2406.18665, first submitted June 26, 2024). It trains routers on human preference data to choose between a stronger and a weaker LLM. The authors report that their routers cut costs by more than 2x in certain cases without compromising response quality, and that the routers kept working when the strong and weak models were swapped at test time. Those results are from 2024-era model pairs. The finding that carries over is that a cheap learned router can capture much of the gap. The exact ratio won't carry over.
Commercial example: OpenRouter's Auto Router. According to OpenRouter's docs, the openrouter/auto model uses "a lightweight classifier" to assign each prompt one of about 30 task types. It then ranks models by how much OpenRouter users spent on that task type over the previous seven days, applies your chosen cost_tier (low through max), and picks a primary model with fallbacks. OpenRouter says there's no extra fee for the router itself: you pay the selected model's rate. You can limit candidates with allowed_models (wildcards such as anthropic/*) or excluded_models, and the response's model field shows which model answered. Note what "powered by the market" means here: the ranking reflects what other customers pay for, which isn't necessarily what's best for your task.
Trade-offs: routers add a decision step (latency, plus a small cost) and can misroute. A misrouted hard request gets a confidently wrong cheap answer, and nothing in the system tells you.
Pattern 2: a cascade with escalation
A cascade sends everything to the cheap model first, then checks the output. The check might be a schema validation, unit tests, a confidence score, or a second model acting as a judge. Anything that fails goes up a tier.
Cascades are attractive when you can verify output mechanically: code that must compile, JSON that must validate, an extraction that must match a regex. They're risky when quality is subjective, because the check itself becomes the weak point. They also cost more than a perfect router, because escalated requests are paid for twice.
Pattern 3: planner/executor, advisors and subagents
The newest pattern is built into the agent tools themselves.
Anthropic's advisor tool (beta). Anthropic's docs describe it this way: "The advisor tool lets a faster, lower-cost executor model consult a higher-intelligence advisor model mid-generation for strategic guidance." The executor (for example, Sonnet) decides when to call the advisor (for example, Opus), and it all happens inside one /v1/messages request. Key details from the docs (re-checked October 9, 2026):
- It needs the beta header
advisor-tool-2026-03-01and a tool of typeadvisor_20260301with a requiredmodel. - The advisor "reads the full conversation". Its tokens are "billed at the advisor model's rates" and don't appear in the top-level
usagetotals. Checkusage.iterations. - The advisor must be Claude Sonnet 4.6 or a more capable model, and at least as capable as the executor. Opus 5, Opus 5.5, Sonnet 5.5, Fable 5/5.1 and Mythos 5/5.1 executors pair only with advisors that return encrypted advice. An unsupported pair returns a 400 error, so check the compatibility table in the docs for your exact pair.
- Newer advisors (Opus 5, Opus 5.5, Sonnet 5.5, Haiku 5.5, Fable and Mythos models) return an encrypted
advisor_redacted_result, so your code can't read the advice. Older advisors such as Opus 4.8 return plain text. - The tool definition's
max_tokens(minimum 1,024) caps each advisor call. The request's ownmax_tokensdoesn't.
Claude Code. Anthropic's coding agent has two built-in routing options. The opusplan model setting "uses opus during plan mode, then switches to sonnet for execution". Subagents take a model field (haiku, sonnet, opus, fable, a full model ID, or inherit). The docs suggest defining your own Explore subagent with model: haiku "to run exploration on a lower-cost model". Anthropic's Haiku 5.5 launch page recommends that kind of pairing: Haiku 5.5 as a subagent alongside Opus 5.5 or Sonnet 5.5. If you're building your own agent, our Claude Agent SDK guide shows where subagents fit.
Managed Agents. Anthropic's August 7, 2026 release notes added "advisors" to the multiagent roster in Claude Managed Agents.
What vendors don't offer (yet)
- Anthropic's
fallbacksparameter isn't a cost router. Its docs say "only a safety classifier decline triggers fallback". Rate limits, overloads and server errors come back as-is. It's in beta, works on the Claude API only, and Haiku 5.5 has no server-side fallback. - OpenAI's API has no automatic tier selection for GPT-6. OpenAI's "Using GPT-6" guide recommends Astra for "the most demanding" work, GPT-6.1 Sol for "balanced speed, cost, and intelligence" and Luna for high-volume tasks, but it describes no automatic selection between them. The only routing it mentions is the older
gpt-5.6alias, which points togpt-5.6-sol. You build the router yourself. - Effort is routing inside a model. Both vendors let you dial reasoning up or down per request. On Anthropic, effort runs from
lowtomax, and Opus 5.5 and Haiku 5.5 default tomedium. On OpenAI, GPT-6 Luna acceptsnonebut GPT-6.1 Sol doesn't. Dropping effort on a mid-tier model is sometimes a better move than dropping to a smaller model, because thinking tokens are billed as output. Our reasoning models explainer covers why.
The cost math (illustrative)
These examples use verified list prices. The traffic mix, request sizes and escalation rates are our assumptions, chosen to show the arithmetic. They aren't measurements.
Assume 1 million requests a month, each with 2,000 input tokens and 500 output tokens.
| Strategy | Calculation | Monthly cost |
|---|---|---|
| Everything on Claude Opus 5.5 ($4/$20) | 1M × (2,000×$4 + 500×$20) ÷ 1M | $18,000 |
| Everything on Claude Sonnet 5.5 ($2/$10) | 1M × (2,000×$2 + 500×$10) ÷ 1M | $9,000 |
| Everything on Claude Haiku 5.5 ($0.10/$0.50) | 1M × (2,000×$0.10 + 500×$0.50) ÷ 1M | $450 |
| Router: 70% Haiku, 25% Sonnet, 5% Opus | $315 + $2,250 + $900 | $3,465 |
| Plus a Haiku 5.5 classifier on every request (2,000 in, 10 out) | 1M × (2,000×$0.10 + 10×$0.50) ÷ 1M | +$205 → $3,670 |
| Cascade: all start on Haiku; 30% re-run on Sonnet; 5% re-run on Opus | $450 + $2,700 + $900 | $4,050 |
| OpenAI equivalent router: 70% Luna, 25% GPT-6.1 Sol, 5% Astra | $315 + $2,250 + $2,250 | $4,815 |
Under these assumptions, routing cuts an all-Opus bill by about 80% and an all-Sonnet bill by about 60%. The cascade costs about 10% more than the router because escalated requests are paid twice, but it doesn't depend on predicting difficulty in advance. The OpenAI mix costs more mainly because its top tier (Astra, $10/$50) is 2.5 times Opus 5.5's price. The real question is how much quality you lose when the router or check gets it wrong, and the table doesn't capture that.
Advisor example (illustrative): three Opus 5.5 advisor calls in a task with a 50K-token conversation, each producing 1,500 tokens, cost about 3 × (50,000 × $4 + 1,500 × $20) ÷ 1M = $0.69, before any caching. Because the advisor reads the full conversation each time, advisor costs grow with context length. Cap max_uses and keep contexts lean.
Two more traps. Haiku 5.5's price rises five-fold once a prompt passes 100K tokens, so a router that sends long documents to "the cheap model" may be less cheap than it looks. And Anthropic's newer models use a tokenizer that produces about 30% more tokens for the same text. Both are covered in why token pricing isn't the story.
How to start
- Log first. Record the input and output tokens and outcome for a week of traffic, and find which endpoints actually need a frontier model.
- Start with rules. Route by endpoint or task type before you train anything.
- Build an eval set of real requests labeled with acceptable answers, and test each tier against it. Our Claude lineup guide lists what each tier costs.
- Add verification where you can, and cascade only where checks are mechanical.
- Consider an open model as the cheap tier if your volume is high. We compare the economics in frontier APIs vs open models.
- Re-run your evals when models change. Router decisions go stale with every release, and Anthropic has released nine Claude models, including two access-restricted Mythos models, on seven dates since June 9, 2026.
What's confirmed and what isn't
- Confirmed from docs: advisor tool behavior and billing,
fallbacksscope, Claude Codeopusplanand the subagentmodelfield, OpenRouter Auto Router mechanics and pricing, and all list prices. - Research result: RouteLLM's "over 2 times" cost reduction in certain cases, on 2024 model pairs.
- Assumptions: every traffic mix and escalation rate in the cost table.
About this storyBased on the sources linked below. Editorial standards




