Distillation and Small Models - Complete Deep Dive
Prerequisites: Fine-Tuning, Model Selection and Routing, Evals Used in: Inference Serving, GPU Capacity and Cost Engineering, Latency Engineering
What is Distillation?
Distillation is using a large capable model to create the training data for a small one. You run the expensive model over inputs drawn from your actual task, keep the outputs that are good, and train a small model on those pairs. The small model does not learn everything the large one knows. It learns the slice you asked about, which for a narrow task is the only slice that was ever going to matter.
The economic insight underneath is worth stating plainly: a large model is often the best way to build a small one. You are converting a recurring per-request cost into a one-time training cost. The expensive model runs once per training example instead of once per user request, forever.
This is the last stage of the track for a reason, and the ordering is not decorative. Prompt first. Retrieve for knowledge. Build evals so you can tell whether anything helped. Fine-tune when the gap is behavioural. Distil when you have a working expensive system, a quality bar you have already measured, and a unit-cost or latency problem that the cheap fixes did not close. Distilling before you have an eval set is the same mistake as fine-tuning before you have one, with an extra model to maintain.
Real-world analogy: A senior engineer who has debugged a class of incident hundreds of times writes the runbook. A new hire follows the runbook and resolves those incidents about as well, far faster, at a fraction of the salary β and falls over on the first incident the runbook does not cover. That is distillation exactly: the expertβs judgement compressed into a narrow, cheap, fast, brittle artifact. The compression is a feature as long as you know the boundary.
The Teacher and Student Flow
flowchart LR
P[Task inputs sampled from real traffic] --> T[Teacher - large capable model]
T --> O[Candidate outputs]
O --> F[Filter - assertions and self consistency and judge]
F --> D[Curated student training set]
D --> S[Train the small student model]
S --> E[Evaluate against your own eval set]
E --> R[Serve the student on the hot path]
R --> B[Escalate hard cases back to the teacher]
classDef client fill:#f97316,stroke:#c2410c,color:#fff
classDef edge fill:#6cf,stroke:#333,color:#000
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef async fill:#b4f,stroke:#333,color:#000
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class P client
class T async
class O,D data
class F edge
class S async
class E service
class R service
class B async
Two details in that diagram carry most of the value.
The filter step is not optional. Teacher output is good, not perfect, and every error you pass through becomes a label the student learns as truth. Filter with whatever you have: schema and assertion checks, agreement across several teacher samples on the same input, a validated judge, and human review of a sample. Dropping the bottom slice of teacher outputs costs you a little data and saves you a lot of inherited noise. The generation and filtering machinery is the same pipeline described in Synthetic Data Generation.
The escalation path is the safety valve. A student plus teacher fallback usually beats a student alone. Route confidently-handled traffic to the small model and send the rest up, using the routing logic in Model Selection and Routing. You keep most of the cost saving and cap the worst-case quality loss.
Distillation Variants, Conceptually
Two broad families, distinguished by how much of the teacher you can see.
Training on the teacherβs outputs. You keep the generated text and train the student on input-output pairs like any supervised fine-tune. This works with any teacher you can call, including a hosted API, which is why it is what product teams actually do. The student sees the answer the teacher committed to and nothing about the alternatives it considered.
Training on the teacherβs output distributions. Where you have access to the teacherβs full next-token probabilities, the student can be trained to match that distribution rather than just the single chosen token. The intuition is that the distribution carries more information than the sample: it says not only βthis wordβ but βand these were the near-misses,β which is a richer target. This needs deeper access than a typical hosted endpoint provides, so in practice it belongs to teams running open-weight teachers themselves.
Variants exist that sit between these β matching intermediate representations, or training on teacher-generated reasoning traces rather than only final answers. The honest summary is that the family is broad and the details are an active area; the product-relevant fact is that output-based distillation is available to everyone and usually sufficient, while distribution-based methods need model access most teams do not have.
Where It Works and Where It Fails
Distillation trades generality for efficiency. That trade is excellent when you do not need generality and terrible when you do.
Works well when the task is narrow and its distribution is stable: a fixed input shape, a bounded output space, and inputs that look broadly like each other. Classification, extraction, routing, reranking, moderation, canned rewrites, and short constrained generation all qualify. The student only has to cover the region you sampled, and you can sample that region thoroughly.
Works badly for open-ended generality. If your product is a general assistant handling arbitrary requests, there is no task distribution to sample β the distribution is βeverything.β A student trained on whatever you happened to collect will be competent on those cases and noticeably worse everywhere else, and βeverywhere elseβ is where your users live. Long multi-step reasoning, novel tool combinations, and anything requiring broad world knowledge degrade first.
The diagnostic question is simple: can you describe your taskβs input distribution well enough to sample it? If yes, distillation is on the table. If the answer is βusers can ask anything,β it is not.
Task-Specific Small Models Are a Production Pattern
Worth saying separately, because it gets treated as a downgrade when it is often the right architecture. Plenty of work inside an AI product is not open-ended generation at all, and for those jobs a small specialised model beats a large general one on cost, on latency, and frequently on accuracy β because a narrow model optimised for one decision has no generality to be distracted by.
| Job | Why a small model fits | Payoff |
|---|---|---|
| Classification - intent labels topic tags sentiment | Fixed label set and short inputs | Cheapest tier by a wide margin and fast enough to run inline |
| Routing - decide which model or tool handles a request | A routing decision must be much cheaper than what it routes to | A router that costs as much as the work defeats its own purpose |
| Extraction - pull structured fields from documents | Bounded output schema with a real ground truth | Directly measurable with reference metrics rather than a judge |
| Reranking - order retrieved candidates by relevance | Scores many pairs per request so per-call cost dominates | Large quality gain in retrieval at a small latency budget |
| Moderation and safety screening - flag risky input or output | Narrow decision that must run on every request | Feasible to apply to all traffic instead of a sample |
Notice that several of these run per candidate rather than per request, which is why model size matters so much more here than on the main generation path. A reranker scoring dozens of candidates per query multiplies its own cost by the candidate count. Latency implications are covered in Latency Engineering.
The Other Size Levers - Quantization and Pruning
Distillation makes a different, smaller model. Quantization and pruning shrink the model you already have. They compose with distillation rather than competing with it.
Quantization stores weights at lower numerical precision. Lower precision means a smaller memory footprint, which means more of the model fits in less memory and more requests fit alongside it β throughput goes up for reasons that are mostly about memory bandwidth rather than arithmetic. Quality cost scales with how far you go: moderate reduction is often close to free in practice, and aggressive reduction starts showing up as degraded output, typically on the hardest inputs first, which is exactly where your eval set needs to be looking.
The practical point that decides architecture: memory footprint usually determines what hardware you need, and hardware class determines cost. A model that fits on a smaller accelerator, or that leaves room for a larger batch on the same one, changes your bill more than most prompt optimisation ever will. See GPU Capacity and Cost Engineering.
Pruning removes weights or whole structural units judged unimportant, then usually retrains to recover. It is real and it is less commonly the first tool a product team reaches for, because the gains are harder to obtain and the serving stack support is thinner than for quantization.
| Technique | Mechanism | What it costs you |
|---|---|---|
| Distillation | Train a small model on a large modelβs outputs for your task | Generality outside the sampled distribution plus a training pipeline and an artifact to own |
| Quantization - moderate | Store weights at lower precision | Usually little quality on typical inputs but it must be verified on your hard cases |
| Quantization - aggressive | Push precision lower still | Visible degradation that shows up first on the hardest inputs |
| Pruning | Remove unimportant weights or units then retrain | Effort and tooling support plus quality loss if pruned too far |
| Smaller base model outright | Just pick a smaller model and prompt it | Capability you may actually need - measure before assuming |
The right order for a cost problem is usually: try a smaller off-the-shelf model with a good prompt first, because it costs an afternoon; quantize next, because it needs no training data; distil last, because it needs a pipeline and an artifact. That ordering is the opposite of how interesting the techniques are.
Licensing and Terms of Service
A caution that belongs in the design review, not in a footnote. Using one providerβs model outputs to train a model β particularly one that could compete with theirs β may be contractually restricted. Terms differ by provider, change over time, and distinguish between usage patterns in ways that are not obvious from the outside. Open-weight models come with their own licences, some permissive about derived models and some not.
The engineering guidance is short: check the actual terms for the teacher you plan to use, before you build the pipeline, and get the answer from someone whose job it is to read them. Do not assume, and do not reason from what other teams appear to be doing. This is a cheap check early and an expensive surprise late.
Evaluation - Only Your Eval Set Counts
There is exactly one question that matters: does the student clear your quality bar on your eval set, on your task?
Public benchmark comparisons are not that question. A small modelβs benchmark averages tell you about general capability across tasks you are not running, on data whose relationship to your inputs is unknown. A student that trails badly on general benchmarks can be the better production choice for your extraction job, and a student that looks strong on them can fail on the one input pattern your users send most.
What a defensible evaluation looks like:
- Teacher as the baseline, scored on the same frozen eval set. That is the quality you are trading away from, so it is the number the student is measured against.
- A prompted small model as the second baseline. If a small off-the-shelf model with a decent prompt matches your student, the distillation pipeline bought you nothing but maintenance.
- Per-category scores, not one average. Distillation degrades unevenly, and an average hides a cliff on one input type. The error-analysis clustering in Evals is how you find it.
- Cost and latency measured alongside quality, because the entire point was the trade. Quantify both sides or you cannot make the call.
- A defined acceptable-loss threshold, agreed before you look at results. Deciding afterwards that the drop is tolerable is not a decision, it is a rationalisation.
- The escalation rate tracked in production. If the student escalates most traffic to the teacher, your savings evaporated and nobody noticed.
Bad to Good to Great
Bad - swap in a small model and hope
Costs are too high, so you point the code at a smaller model, skim a few outputs, and ship. Quality drops in a way that is invisible in aggregate and obvious on the hard inputs β which are the ones that mattered β and you find out from users. You also cannot go back cleanly, because you never measured what you had.
Good - distil on teacher outputs and measure the average
You sample real inputs, generate teacher outputs, train a student, and confirm the average eval score is close enough. Genuine engineering. Two gaps remain: the average is hiding whichever category fell off a cliff, and you never checked whether a plainly prompted small model would have matched the student without a pipeline to maintain.
Great - measured trade, filtered data, escalation path, tracked in production
- Cheap levers first β a smaller off-the-shelf model with a good prompt, then quantization, before building any pipeline.
- Task distribution sampled from real traffic, with the hard and rare cases deliberately over-represented.
- Teacher outputs filtered by assertions, cross-sample agreement, a validated judge, and human review of a sample, so the student does not inherit teacher errors as labels.
- Teacher and prompted-small-model baselines both scored on a frozen eval set, with per-category breakdowns.
- An acceptable quality loss agreed in advance, and cost and latency quantified against it.
- Student plus teacher escalation in serving, so worst-case quality is capped, with the escalation rate monitored as a cost metric.
- Licence and terms checked for the teacher before the pipeline exists.
- A re-distillation trigger defined β drift, a teacher upgrade, or an eval regression β because the student is a snapshot of a task that keeps moving.
When to Use
β Distil or go small when:
- A narrow high-volume task dominates your inference bill and the input distribution is describable
- Latency on the hot path matters more than breadth of capability
- The job is classification, routing, extraction, reranking, or moderation, where specialised models are simply the better fit
- You already have a working expensive system and a measured quality bar to trade against
- You need to run a model where a large one cannot go, such as on constrained hardware or fully in your own environment
β Do not distil when:
- You have no eval set, because you will be unable to tell what the compression cost you
- The product is open-ended and the input distribution is effectively unbounded
- You have not yet tried a smaller off-the-shelf model with a serious prompt, which is an afternoon rather than a pipeline
- The teacherβs terms of service restrict training on its output for your intended use
- The task needs broad world knowledge or long multi-step reasoning, which is what degrades first
- Nobody will own the pipeline, the artifact, and its eventual re-distillation
Common Interview Questions
Q1: Why would you use a large model to build a small one rather than just using the small one?
Because the expensive model is better at producing the training data than any other source available, and you only pay for it once per example instead of once per request forever. A small model prompted from scratch has to infer the task from your instructions; a small model trained on a few thousand high-quality teacher outputs on your actual input distribution has seen the task done well repeatedly. That converts a recurring per-request cost into a one-time training cost, and on a narrow task it buys most of the teacherβs quality at a fraction of cost and latency. The catch is that it only holds inside the distribution you sampled β outside it, the student is worse than the small model looked in your demo.
Q2: When is distillation the wrong tool?
When there is no describable task distribution to sample. If the product is an open-ended assistant where users can ask anything, the student will be competent on whatever you collected and noticeably worse everywhere else, and everywhere else is most of the traffic. Broad world knowledge and long multi-step reasoning degrade first. It is also wrong before you have an eval set, since you cannot measure what the compression cost; wrong before you have tried a smaller off-the-shelf model with a real prompt, because that is an afternoon and this is a pipeline; and wrong if the teacherβs terms restrict training on its output for your use.
Q3: How do quantization and distillation differ, and which do you try first?
Distillation trains a different, smaller model on your task. Quantization keeps the same model and stores its weights at lower precision, cutting memory footprint and raising throughput largely through memory bandwidth rather than arithmetic. Quantization is the cheaper experiment because it needs no training data and no artifact β you try it and measure. So the order for a cost problem is a smaller off-the-shelf model with a good prompt, then quantization, then distillation. The reason memory footprint gets so much attention is that it usually decides which hardware class you need and how large a batch fits, and that decides the bill far more than prompt micro-optimisation does.
Q4: How do you evaluate a distilled model?
Against my own eval set on my own task, with the teacher as the baseline, because the teacherβs score is the quality I am trading away from. I add a second baseline β a plainly prompted small model β since if that matches the student, the pipeline bought nothing but maintenance. I read per-category scores rather than one average, because distillation degrades unevenly and an average hides a cliff on one input type. I agree the acceptable quality loss before seeing results, so the decision is a decision and not a rationalisation. And I measure cost and latency alongside quality, since the trade was the entire point. Public benchmark numbers are not part of this β they describe tasks I am not running.
Q5: Your distilled model is slightly worse than the teacher. How do you ship it responsibly?
With an escalation path and a monitored boundary. Serve the student on traffic it handles confidently and route the rest to the teacher, which keeps most of the cost saving while capping worst-case quality. Then instrument the boundary: track the escalation rate as a cost metric, because a student that escalates most traffic has quietly deleted the savings, and track per-category quality in production rather than relying on the pre-launch average. Roll out behind a flag on a slice of traffic with a defined rollback and the previous path kept warm. And define the re-distillation trigger up front β drift, a teacher upgrade, or an eval regression β since the student is a snapshot of a task distribution that keeps moving.
Build it in code: Agentic AI Course · Fundamentals: Core Concepts