Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 17 min read

Distillation and Small Models - Complete Deep Dive

Stage 6 - Customization Lesson 27 of 36

Prerequisites: Fine-Tuning, Model Selection and Routing, Evals Used in: Inference Serving, GPU Capacity and Cost Engineering, Latency Engineering


What is Distillation?

Distillation is using a large capable model to create the training data for a small one. You run the expensive model over inputs drawn from your actual task, keep the outputs that are good, and train a small model on those pairs. The small model does not learn everything the large one knows. It learns the slice you asked about, which for a narrow task is the only slice that was ever going to matter.

The economic insight underneath is worth stating plainly: a large model is often the best way to build a small one. You are converting a recurring per-request cost into a one-time training cost. The expensive model runs once per training example instead of once per user request, forever.

This is the last stage of the track for a reason, and the ordering is not decorative. Prompt first. Retrieve for knowledge. Build evals so you can tell whether anything helped. Fine-tune when the gap is behavioural. Distil when you have a working expensive system, a quality bar you have already measured, and a unit-cost or latency problem that the cheap fixes did not close. Distilling before you have an eval set is the same mistake as fine-tuning before you have one, with an extra model to maintain.

Real-world analogy: A senior engineer who has debugged a class of incident hundreds of times writes the runbook. A new hire follows the runbook and resolves those incidents about as well, far faster, at a fraction of the salary β€” and falls over on the first incident the runbook does not cover. That is distillation exactly: the expert’s judgement compressed into a narrow, cheap, fast, brittle artifact. The compression is a feature as long as you know the boundary.


The Teacher and Student Flow

flowchart LR
    P[Task inputs sampled from real traffic] --> T[Teacher - large capable model]
    T --> O[Candidate outputs]
    O --> F[Filter - assertions and self consistency and judge]
    F --> D[Curated student training set]
    D --> S[Train the small student model]
    S --> E[Evaluate against your own eval set]
    E --> R[Serve the student on the hot path]
    R --> B[Escalate hard cases back to the teacher]

    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef edge fill:#6cf,stroke:#333,color:#000
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef async fill:#b4f,stroke:#333,color:#000
    classDef data fill:#fbbf24,stroke:#92400e,color:#000

    class P client
    class T async
    class O,D data
    class F edge
    class S async
    class E service
    class R service
    class B async

Two details in that diagram carry most of the value.

The filter step is not optional. Teacher output is good, not perfect, and every error you pass through becomes a label the student learns as truth. Filter with whatever you have: schema and assertion checks, agreement across several teacher samples on the same input, a validated judge, and human review of a sample. Dropping the bottom slice of teacher outputs costs you a little data and saves you a lot of inherited noise. The generation and filtering machinery is the same pipeline described in Synthetic Data Generation.

The escalation path is the safety valve. A student plus teacher fallback usually beats a student alone. Route confidently-handled traffic to the small model and send the rest up, using the routing logic in Model Selection and Routing. You keep most of the cost saving and cap the worst-case quality loss.


Distillation Variants, Conceptually

Two broad families, distinguished by how much of the teacher you can see.

Training on the teacher’s outputs. You keep the generated text and train the student on input-output pairs like any supervised fine-tune. This works with any teacher you can call, including a hosted API, which is why it is what product teams actually do. The student sees the answer the teacher committed to and nothing about the alternatives it considered.

Training on the teacher’s output distributions. Where you have access to the teacher’s full next-token probabilities, the student can be trained to match that distribution rather than just the single chosen token. The intuition is that the distribution carries more information than the sample: it says not only β€œthis word” but β€œand these were the near-misses,” which is a richer target. This needs deeper access than a typical hosted endpoint provides, so in practice it belongs to teams running open-weight teachers themselves.

Variants exist that sit between these β€” matching intermediate representations, or training on teacher-generated reasoning traces rather than only final answers. The honest summary is that the family is broad and the details are an active area; the product-relevant fact is that output-based distillation is available to everyone and usually sufficient, while distribution-based methods need model access most teams do not have.


Where It Works and Where It Fails

Distillation trades generality for efficiency. That trade is excellent when you do not need generality and terrible when you do.

Works well when the task is narrow and its distribution is stable: a fixed input shape, a bounded output space, and inputs that look broadly like each other. Classification, extraction, routing, reranking, moderation, canned rewrites, and short constrained generation all qualify. The student only has to cover the region you sampled, and you can sample that region thoroughly.

Works badly for open-ended generality. If your product is a general assistant handling arbitrary requests, there is no task distribution to sample β€” the distribution is β€œeverything.” A student trained on whatever you happened to collect will be competent on those cases and noticeably worse everywhere else, and β€œeverywhere else” is where your users live. Long multi-step reasoning, novel tool combinations, and anything requiring broad world knowledge degrade first.

The diagnostic question is simple: can you describe your task’s input distribution well enough to sample it? If yes, distillation is on the table. If the answer is β€œusers can ask anything,” it is not.


Task-Specific Small Models Are a Production Pattern

Worth saying separately, because it gets treated as a downgrade when it is often the right architecture. Plenty of work inside an AI product is not open-ended generation at all, and for those jobs a small specialised model beats a large general one on cost, on latency, and frequently on accuracy β€” because a narrow model optimised for one decision has no generality to be distracted by.

Job Why a small model fits Payoff
Classification - intent labels topic tags sentiment Fixed label set and short inputs Cheapest tier by a wide margin and fast enough to run inline
Routing - decide which model or tool handles a request A routing decision must be much cheaper than what it routes to A router that costs as much as the work defeats its own purpose
Extraction - pull structured fields from documents Bounded output schema with a real ground truth Directly measurable with reference metrics rather than a judge
Reranking - order retrieved candidates by relevance Scores many pairs per request so per-call cost dominates Large quality gain in retrieval at a small latency budget
Moderation and safety screening - flag risky input or output Narrow decision that must run on every request Feasible to apply to all traffic instead of a sample

Notice that several of these run per candidate rather than per request, which is why model size matters so much more here than on the main generation path. A reranker scoring dozens of candidates per query multiplies its own cost by the candidate count. Latency implications are covered in Latency Engineering.


The Other Size Levers - Quantization and Pruning

Distillation makes a different, smaller model. Quantization and pruning shrink the model you already have. They compose with distillation rather than competing with it.

Quantization stores weights at lower numerical precision. Lower precision means a smaller memory footprint, which means more of the model fits in less memory and more requests fit alongside it β€” throughput goes up for reasons that are mostly about memory bandwidth rather than arithmetic. Quality cost scales with how far you go: moderate reduction is often close to free in practice, and aggressive reduction starts showing up as degraded output, typically on the hardest inputs first, which is exactly where your eval set needs to be looking.

The practical point that decides architecture: memory footprint usually determines what hardware you need, and hardware class determines cost. A model that fits on a smaller accelerator, or that leaves room for a larger batch on the same one, changes your bill more than most prompt optimisation ever will. See GPU Capacity and Cost Engineering.

Pruning removes weights or whole structural units judged unimportant, then usually retrains to recover. It is real and it is less commonly the first tool a product team reaches for, because the gains are harder to obtain and the serving stack support is thinner than for quantization.

Technique Mechanism What it costs you
Distillation Train a small model on a large model’s outputs for your task Generality outside the sampled distribution plus a training pipeline and an artifact to own
Quantization - moderate Store weights at lower precision Usually little quality on typical inputs but it must be verified on your hard cases
Quantization - aggressive Push precision lower still Visible degradation that shows up first on the hardest inputs
Pruning Remove unimportant weights or units then retrain Effort and tooling support plus quality loss if pruned too far
Smaller base model outright Just pick a smaller model and prompt it Capability you may actually need - measure before assuming

The right order for a cost problem is usually: try a smaller off-the-shelf model with a good prompt first, because it costs an afternoon; quantize next, because it needs no training data; distil last, because it needs a pipeline and an artifact. That ordering is the opposite of how interesting the techniques are.


Licensing and Terms of Service

A caution that belongs in the design review, not in a footnote. Using one provider’s model outputs to train a model β€” particularly one that could compete with theirs β€” may be contractually restricted. Terms differ by provider, change over time, and distinguish between usage patterns in ways that are not obvious from the outside. Open-weight models come with their own licences, some permissive about derived models and some not.

The engineering guidance is short: check the actual terms for the teacher you plan to use, before you build the pipeline, and get the answer from someone whose job it is to read them. Do not assume, and do not reason from what other teams appear to be doing. This is a cheap check early and an expensive surprise late.


Evaluation - Only Your Eval Set Counts

There is exactly one question that matters: does the student clear your quality bar on your eval set, on your task?

Public benchmark comparisons are not that question. A small model’s benchmark averages tell you about general capability across tasks you are not running, on data whose relationship to your inputs is unknown. A student that trails badly on general benchmarks can be the better production choice for your extraction job, and a student that looks strong on them can fail on the one input pattern your users send most.

What a defensible evaluation looks like:

  1. Teacher as the baseline, scored on the same frozen eval set. That is the quality you are trading away from, so it is the number the student is measured against.
  2. A prompted small model as the second baseline. If a small off-the-shelf model with a decent prompt matches your student, the distillation pipeline bought you nothing but maintenance.
  3. Per-category scores, not one average. Distillation degrades unevenly, and an average hides a cliff on one input type. The error-analysis clustering in Evals is how you find it.
  4. Cost and latency measured alongside quality, because the entire point was the trade. Quantify both sides or you cannot make the call.
  5. A defined acceptable-loss threshold, agreed before you look at results. Deciding afterwards that the drop is tolerable is not a decision, it is a rationalisation.
  6. The escalation rate tracked in production. If the student escalates most traffic to the teacher, your savings evaporated and nobody noticed.

Bad to Good to Great

Bad - swap in a small model and hope

Costs are too high, so you point the code at a smaller model, skim a few outputs, and ship. Quality drops in a way that is invisible in aggregate and obvious on the hard inputs β€” which are the ones that mattered β€” and you find out from users. You also cannot go back cleanly, because you never measured what you had.

Good - distil on teacher outputs and measure the average

You sample real inputs, generate teacher outputs, train a student, and confirm the average eval score is close enough. Genuine engineering. Two gaps remain: the average is hiding whichever category fell off a cliff, and you never checked whether a plainly prompted small model would have matched the student without a pipeline to maintain.

Great - measured trade, filtered data, escalation path, tracked in production

  1. Cheap levers first β€” a smaller off-the-shelf model with a good prompt, then quantization, before building any pipeline.
  2. Task distribution sampled from real traffic, with the hard and rare cases deliberately over-represented.
  3. Teacher outputs filtered by assertions, cross-sample agreement, a validated judge, and human review of a sample, so the student does not inherit teacher errors as labels.
  4. Teacher and prompted-small-model baselines both scored on a frozen eval set, with per-category breakdowns.
  5. An acceptable quality loss agreed in advance, and cost and latency quantified against it.
  6. Student plus teacher escalation in serving, so worst-case quality is capped, with the escalation rate monitored as a cost metric.
  7. Licence and terms checked for the teacher before the pipeline exists.
  8. A re-distillation trigger defined β€” drift, a teacher upgrade, or an eval regression β€” because the student is a snapshot of a task that keeps moving.

When to Use

βœ… Distil or go small when:

❌ Do not distil when:


Common Interview Questions

Q1: Why would you use a large model to build a small one rather than just using the small one?

Because the expensive model is better at producing the training data than any other source available, and you only pay for it once per example instead of once per request forever. A small model prompted from scratch has to infer the task from your instructions; a small model trained on a few thousand high-quality teacher outputs on your actual input distribution has seen the task done well repeatedly. That converts a recurring per-request cost into a one-time training cost, and on a narrow task it buys most of the teacher’s quality at a fraction of cost and latency. The catch is that it only holds inside the distribution you sampled β€” outside it, the student is worse than the small model looked in your demo.

Q2: When is distillation the wrong tool?

When there is no describable task distribution to sample. If the product is an open-ended assistant where users can ask anything, the student will be competent on whatever you collected and noticeably worse everywhere else, and everywhere else is most of the traffic. Broad world knowledge and long multi-step reasoning degrade first. It is also wrong before you have an eval set, since you cannot measure what the compression cost; wrong before you have tried a smaller off-the-shelf model with a real prompt, because that is an afternoon and this is a pipeline; and wrong if the teacher’s terms restrict training on its output for your use.

Q3: How do quantization and distillation differ, and which do you try first?

Distillation trains a different, smaller model on your task. Quantization keeps the same model and stores its weights at lower precision, cutting memory footprint and raising throughput largely through memory bandwidth rather than arithmetic. Quantization is the cheaper experiment because it needs no training data and no artifact β€” you try it and measure. So the order for a cost problem is a smaller off-the-shelf model with a good prompt, then quantization, then distillation. The reason memory footprint gets so much attention is that it usually decides which hardware class you need and how large a batch fits, and that decides the bill far more than prompt micro-optimisation does.

Q4: How do you evaluate a distilled model?

Against my own eval set on my own task, with the teacher as the baseline, because the teacher’s score is the quality I am trading away from. I add a second baseline β€” a plainly prompted small model β€” since if that matches the student, the pipeline bought nothing but maintenance. I read per-category scores rather than one average, because distillation degrades unevenly and an average hides a cliff on one input type. I agree the acceptable quality loss before seeing results, so the decision is a decision and not a rationalisation. And I measure cost and latency alongside quality, since the trade was the entire point. Public benchmark numbers are not part of this β€” they describe tasks I am not running.

Q5: Your distilled model is slightly worse than the teacher. How do you ship it responsibly?

With an escalation path and a monitored boundary. Serve the student on traffic it handles confidently and route the rest to the teacher, which keeps most of the cost saving while capping worst-case quality. Then instrument the boundary: track the escalation rate as a cost metric, because a student that escalates most traffic has quietly deleted the savings, and track per-category quality in production rather than relying on the pre-launch average. Roll out behind a flag on a slice of traffic with a defined rollback and the previous path kept warm. And define the re-distillation trigger up front β€” drift, a teacher upgrade, or an eval regression β€” since the student is a snapshot of a task distribution that keeps moving.


Build it in code: Agentic AI Course · Fundamentals: Core Concepts

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access