Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
โฑ๏ธ 17 min read

Model Selection and Routing - Complete Deep Dive

Prerequisites: LLM APIs and SDKs, Tokens and Cost Math, Evals Used in: Latency Engineering, GPU Cost Engineering, Agent Architectures, Deploying AI Features


What is Model Selection and Routing?

Model selection is choosing which model serves a given task. Routing is doing that choice per request at runtime instead of once at design time. Together they are the highest-leverage cost and latency decision in most LLM applications, and the one teams most often make by accident โ€” picking whatever they used in the prototype and never revisiting it.

Real-world analogy: A law firm does not staff every matter with a senior partner. Intake forms go to a paralegal, standard contracts to an associate, and the genuinely novel dispute to the partner. The paralegal escalates when a matter turns out to be harder than the intake suggested. Nobody argues the partner is worse at reviewing intake forms โ€” they are better at it and also twenty times the hourly rate, so using them for it is a budgeting failure, not a quality win. Routing is staffing.


The Axes That Actually Decide It

There are eight. Most teams weigh the first and ignore the rest, which is how you end up with a system that is accurate, beautiful, too slow, and unaffordable.

Axis The question to ask What a wrong answer costs
Task difficulty Is this extraction and formatting, or multi-step reasoning? Oversized model on trivial work, or a small model quietly failing on hard work
Quality bar What failure rate is acceptable, and who sees a failure? Chasing accuracy nobody needs, or shipping errors into a legal document
Latency budget Interactive under a second, or a batch job overnight? Larger models generate more slowly; a P95 breach is a product defect
Cost per call Cost times volume โ€” what is the monthly bill at projected traffic? The frontier model that works in the demo and cannot survive launch
Context length How much input per call, realistically at P95 and not average? Truncated context, silent quality loss, or paying for a long window you never fill
Tool and schema reliability Does it produce valid tool calls and valid schemas consistently? An agent that reasons well and cannot call anything correctly
Residency and control Can this data leave your boundary at all? A compliance incident, which is not a tradeoff you get to make later
Lock-in How much rewriting if this provider degrades, reprices, or deprecates? A migration you cannot execute under time pressure

Two of these are gates rather than dials. Data residency either permits a hosted API or it does not; no amount of quality compensates. And a hard latency budget eliminates whole model classes before quality enters the conversation.


The Biggest Model Is Not a Safe Default

โ€œUse the strongest model availableโ€ feels like the conservative choice. It is not conservative, it is unmeasured โ€” and it ships three defects at once.

It is a cost bug. The price spread between a frontier model and a small one in the same family is routinely an order of magnitude per token, sometimes more. Multiply by volume. A feature that is comfortably profitable on a small model can be structurally unprofitable on a large one, and you will discover this after launch when the bill arrives rather than before.

It is a latency bug. Bigger models generate tokens more slowly. On an interactive surface, the difference between a fast small model and a frontier model is often the difference between a feature that feels responsive and one users abandon mid-response โ€” and users do not file a ticket saying โ€œyour model is too large,โ€ they just leave.

It is an evidence bug. If you never tested a smaller model, you do not know the large one was necessary. You have a working system and no idea where it sits on the cost-quality frontier.

That frontier is the useful mental model: plot quality against cost per call for every candidate on your task. Most models are strictly dominated โ€” something is both better and cheaper. A few sit on the efficient frontier, and your job is choosing a point on it rather than defaulting to the top-right corner.

The observation that makes routing pay is this: most production traffic is easy traffic. Real workloads are heavily skewed โ€” a long tail of hard, novel requests and a fat head of repetitive, formulaic ones. Sizing your entire fleet for the tail means overpaying on the head, which is the majority of your volume. You do not need to guess the split; label a sample of real traffic by difficulty and count.


Routing Patterns

Pattern How it decides Strengths Weaknesses
Static by task type Code maps each call site to a model Trivial to reason about, zero added latency, easy to test Cannot adapt within a task type
Cascade or escalation Small model first, escalate on a failure signal Captures the easy-traffic win with a correctness backstop Escalated requests pay both calls in latency and cost
Classifier router A small model or trained classifier predicts difficulty up front One call per request, no double payment The classifier is itself a component that can be wrong, drift, and need evals

Static routing is where you start, and for many systems it is where you stop. You already know that title generation is easy and root-cause analysis is hard. Encode it: a cheap model for summarization and extraction, a strong one for the reasoning step, a fast one for anything on the userโ€™s critical path. It is a config table, it costs nothing at runtime, and it captures most of the available savings.

Cascades go further, and the design work is entirely in the escalation signal. Options, roughly in order of reliability:

flowchart LR
    A[Incoming request] --> B[Router]
    B -->|Known hard task type| F[Large model]
    B -->|Default path| C[Small fast model]
    C --> D[Check schema and confidence]
    D -->|Passes| E[Response to caller]
    D -->|Fails or abstains| F
    F --> G[Check schema]
    G -->|Passes| E
    G -->|Fails| H[Escalate to human or fail loudly]

    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef edge fill:#6cf,stroke:#333,color:#000
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef async fill:#b4f,stroke:#333,color:#000
    classDef data fill:#fbbf24,stroke:#92400e,color:#000

    class A client
    class B edge
    class C,F service
    class D,G async
    class E,H data

The arithmetic decides whether a cascade is worth it. Escalated requests pay for both calls and both latencies, so the win depends on the escalation rate and the price ratio. At a low escalation rate with a large price gap, a cascade is a clear win. At a high escalation rate, you have built a slower, more expensive system than just calling the strong model โ€” and the escalation rate is an empirical number from your traffic, not an assumption. Instrument it and alert on it: a climbing escalation rate is your earliest signal that inputs have drifted.


Hosted API vs Open-Weights Self-Hosted

ย  Hosted API Open-weights self-hosted
Time to first call Minutes Days to weeks
Cost shape Per token, scales with usage Fixed GPU capacity, scales with provisioning
Cost at low volume Cheap Expensive โ€” idle GPUs bill anyway
Cost at high steady volume Can dominate the bill Often much cheaper per token
Data boundary Leaves your perimeter Stays inside it
Version control Provider decides deprecations You decide, indefinitely
Fine-tuning freedom Limited to what the provider offers Full weights, any technique
Operational burden Providerโ€™s problem Yours: capacity, upgrades, on-call, autoscaling

Self-hosting genuinely wins in three cases, and it is worth being precise because the decision is usually made on vibes:

  1. Volume. A high, steady, predictable request rate is where per-token pricing loses to amortized GPU cost. The crossover is real and computable from your own numbers โ€” do that arithmetic before arguing about it (GPU Capacity and Cost Engineering).
  2. Privacy and residency. When data cannot cross a boundary, this is not an optimization, it is the only option.
  3. Tuning and control. Full weights mean any fine-tuning or distillation technique, plus an upgrade schedule nobody else can change under you.

The honest cost side: you now own inference serving. Continuous batching, KV-cache management, and GPU utilization become your problem (Inference Serving), along with capacity planning for peaks, model upgrades, quantization tradeoffs, and a new on-call surface. Idle GPUs bill at full rate, so spiky traffic is the worst possible fit. Teams routinely underestimate this by an order of magnitude and end up with a worse-performing endpoint than the API they left. A common middle path is worth naming: hosted APIs for the low-volume, high-difficulty tail and a self-hosted small model for the high-volume head, which puts your fixed capacity exactly where utilization is predictable.


Evaluating a Candidate Model on Your Task

Public leaderboards are a starting filter for which models to try. They are weak evidence for which model to ship, for four reasons:

The procedure that does produce evidence, in order:

  1. Build a golden set from real traffic โ€” 50 to a few hundred labelled cases, weighted toward boundary and failure cases (Evals)
  2. Run every candidate on the same prompt for a fair baseline, then also run a lightly per-model tuned prompt, because prompts are not portable across families and a model can lose on a prompt written for a competitor
  3. Measure four things, not one: quality score, P95 latency, cost per call, and format-compliance rate โ€” a model that is slightly less accurate but never breaks your schema can be the better production choice
  4. Test the failure modes you care about: long inputs at your real P95, adversarial inputs, tool-call correctness, and behavior on inputs it should refuse or abstain on
  5. Shadow real traffic before switching. Send a slice to the candidate without using its output, and compare offline. This catches distribution surprises your golden set missed
  6. Roll out progressively, with the old model one config change away (Deploying AI Features)

A Thin Abstraction and a Pinning Policy

You will change models. Provider pricing moves, capable models ship, a cheaper option becomes good enough, a provider deprecates what you depend on. Design so that switching is a config change and an eval run, not a refactor.

Thin is the operative word. The failure mode on both sides is well known: hard-coding one providerโ€™s SDK through your handlers makes a swap a rewrite, while building a grand abstraction over every providerโ€™s every feature makes you maintain a framework, and the abstraction leaks anyway the moment you need something provider-specific.

A boundary that holds in practice:

class Completion:
    text: str
    input_tokens: int
    output_tokens: int
    model_id: str          # exact pinned version that served this call
    finish_reason: str

class ModelClient(Protocol):
    def complete(
        self, messages: list[Message], schema: dict | None = None, max_tokens: int = 1024
    ) -> Completion: ...

MODEL_ROUTES = {
    "ticket.triage":    "<small-model>@<pinned-snapshot>",   # placeholders - never an alias
    "ticket.rootcause": "<reasoning-model>@<pinned-snapshot>",
}

Keep the interface to what you actually use across providers โ€” messages in, text plus token counts plus a schema option out. Put provider-specific features behind explicit escape hatches rather than pretending they generalize. And keep the route table as data, so changing which model serves a call site is a config diff someone can review and revert.

Pin versions. Calling a floating alias like latest means a provider-side update can change your behavior with no deploy on your side โ€” quality, formatting, latency, and cost all shift under you, and your incident timeline shows no change because there was none in your repo. Pin an explicit snapshot, record the exact model_id on every request, and treat a version bump as a code change: run the eval suite, diff the scores, canary, then expand. Track deprecation notices, because pinned versions do eventually retire, and a forced upgrade you have two weeks to absorb is much easier when your eval suite already exists.


When to Use

โœ… Introduce multi-model routing when:

โŒ Stay on one model when:


Common Interview Questions

Q1: Why is defaulting to the strongest available model a bad engineering decision?

Because it optimizes one axis and silently fails three others. Cost: the spread between a frontier and a small model in the same family is routinely an order of magnitude per token, so a feature that is profitable on a small model can be structurally unprofitable on a large one. Latency: bigger models generate more slowly, and on an interactive surface that can be the difference between a feature people use and one they abandon. Evidence: if you never tested a smaller model you do not know the large one was necessary, so you have no idea where you sit on the cost-quality frontier. The defensible version of the decision is to establish the quality bar, find the cheapest model that clears it on your eval set, and reserve the expensive model for traffic that actually needs it.

Q2: Design a cascade router. What escalates, and how do you know it is worth it?

Route to a small model first. Escalate on an objective signal: schema validation failure, an explicit abstention the prompt permits, or a judge score below a threshold. Self-reported confidence is the weakest of these because models are poorly calibrated about their own certainty, so it should be a hint rather than a gate. Worth-it is arithmetic: escalated requests pay both calls in cost and latency, so the win depends on the escalation rate against the price ratio between the two models. Low escalation rate with a wide price gap is a clear win; a high escalation rate means you built something slower and more expensive than just calling the strong model. Both numbers come from your traffic, so instrument the escalation rate and alert on it โ€” a climbing rate is your earliest signal of input drift.

Q3: A model tops the public leaderboards. Why is that weak evidence for your product?

Four reasons. Contamination: public test sets leak into training data, so part of the score may be memorization, and you cannot audit a closed modelโ€™s training set to check. Distribution mismatch: your inputs are not the benchmarkโ€™s โ€” your formats, jargon, languages, and malformed real-world data. Aggregation: a headline number averages over subtasks and can hide that the third-ranked model is best at your specific one. Prompt sensitivity: leaderboard numbers come from the benchmarkโ€™s prompt, and rankings can reorder under yours. Leaderboards are a filter for what to try. The decision comes from your own golden set, measured on quality, P95 latency, cost per call, and format-compliance rate, ideally confirmed by shadowing real traffic before you switch.

Q4: When does self-hosting an open-weights model actually beat a hosted API?

Three cases. High, steady, predictable volume, where amortized GPU cost beats per-token pricing โ€” and the crossover is computable from your own numbers rather than a matter of opinion. Data that cannot leave your boundary, where it is not an optimization but the only legal option. And when you need full control over weights and upgrade timing, for fine-tuning or distillation, or to stop a provider retiring a version under you. The cost you take on is an inference platform: continuous batching, KV-cache and utilization management, capacity for peaks, quantization tradeoffs, and a new on-call surface. Idle GPUs bill at full rate, so spiky traffic is the worst fit. A common compromise is a self-hosted small model for the high-volume predictable head and a hosted API for the low-volume hard tail.

Q5: How do you keep a provider-side model update from silently changing your product?

Pin an explicit model version rather than calling a floating alias, so nothing changes without a deploy on your side. Record the exact model identifier on every request, so a behavior change can be attributed rather than guessed at. Keep the model choice in a config table behind a thin client interface, so switching is a reviewable config diff and a rollback is one change. Treat any version bump as a code change: run the eval suite, diff the scores against the pinned baseline, canary on a small traffic slice, watch quality and cost and latency, then expand. And track deprecation notices, because pinned versions do retire โ€” a forced upgrade on a two-week deadline is survivable only if the eval suite already exists.


Build it in code: Agentic AI Course · Fundamentals: Core Concepts

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access