Fine-Tuning - Complete Deep Dive
Prerequisites: How LLMs Actually Work, Prompt Engineering, Evals Used in: Preference Tuning, Distillation and Small Models, Inference Serving
What is Fine-Tuning?
Fine-tuning is continuing to train an already-trained model on your own examples so it learns a behaviour you could not reliably get from a prompt. You supply input-output pairs that demonstrate the behaviour you want, the training process nudges the weights toward producing those outputs, and you end up with a new model artifact that is yours to version, host, and eventually retire.
Say the uncomfortable part first, because this whole stage of the track depends on it: fine-tuning is the technique you reach for last, and it is the one people reach for first. The most common mistake in applied AI is trying to fine-tune away a knowledge gap. A model that does not know your Q3 pricing does not need training β it needs the pricing document in its context window. Training facts into weights is slower, more expensive, harder to update, impossible to cite, and worse at the job than a retrieval call you can ship this afternoon. When the facts change next quarter you retrain; the retrieval system just re-indexes.
Real-world analogy: Prompting is giving an employee written instructions. Retrieval is handing them the current policy binder. Fine-tuning is sending them on a two-week training course so the behaviour becomes reflex. Courses are great for βalways structure the report this wayβ and useless for βthe price changed yesterdayβ β nobody re-runs onboarding to communicate a price change. The person who insists on a training course for every problem is the person who has not noticed the binder on the shelf.
What Fine-Tuning Changes and What It Does Not
| Fine-tuning reliably changes | Fine-tuning does not reliably do |
|---|---|
| Behaviour and habits - always ask a clarifying question before acting, always abstain when evidence is thin | Inject new facts. Some facts stick, some do not, and you cannot tell which. That is retrievalβs job. |
| Format adherence - a house output shape the model hits every time without being re-told | Keep facts current. Weights are frozen at training time. Your corpus is not. |
| Tone and style - your voice rather than the providerβs default register | Provide citations. There is no source document to point at, so groundedness is unverifiable. |
| Task-specific skill - a narrow job done well by a small model, cheaply and fast | Fix a bad instruction. If the prompt was ambiguous, tuning teaches the model to guess your ambiguity consistently. |
| Implicit conventions - domain shorthand and internal vocabulary you would otherwise spend hundreds of tokens explaining | Add reasoning ability the base model lacks. Tuning shapes a capability; it rarely creates one. |
The pattern in the right column: fine-tuning moves the modelβs disposition, not its information. Treat any proposal that depends on the model memorising a fact as a retrieval proposal in disguise β the mechanics are in RAG End to End.
The Decision Framework
Before anything else, classify the failure. Nearly every fine-tuning project that gets cancelled halfway through was misclassified on day one.
flowchart TD
S[Symptom - output is not what you want] --> A[Read the failures and classify the cause]
A -->|Missing or stale facts| R[Retrieval - ground the answer in documents]
A -->|Instruction ignored or misread| P[Prompt engineering - sharper spec and examples]
A -->|Schema drifts between runs| F[Fine-tune for format adherence]
A -->|House tone never sticks| F
A -->|Narrow task too slow or costly on a large model| F
A -->|Judgement calls feel systematically off| PT[Preference tuning]
R --> E[Re-measure on the same eval set]
P --> E
F --> E
PT --> E
classDef client fill:#f97316,stroke:#c2410c,color:#fff
classDef edge fill:#6cf,stroke:#333,color:#000
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class S client
class A edge
class R,P service
class F,PT data
class E service
| Symptom | Real cause | Right tool | Why not fine-tuning |
|---|---|---|---|
| Model does not know your products, policies, or docs | Information is not in context | Retrieval | Facts in weights go stale and cannot be cited |
| Answers cite nothing and you need provenance | No source to attribute | Retrieval | A tuned model has no document to point at |
| Model ignores a rule you stated once in a long prompt | Instruction buried or vague | Prompting - restructure, move rules up, add examples | You would be training in your own ambiguity |
| Output shape is right 90 percent of the time | Format under-specified | Structured outputs first, then tuning if it still drifts | Schema-constrained decoding is cheaper and exact |
| Every answer sounds like a generic assistant | Style is not learnable from a paragraph of description | Fine-tuning | This is the job it is good at |
| One narrow high-volume task is dominating cost | Large model doing small work | Fine-tuning a small model | Nothing else closes a cost gap this large |
| Long few-shot block in every request | Examples are paying rent per call | Fine-tuning to absorb the examples | Prompt length is a recurring cost and a latency tax |
| Refusals and hedging feel miscalibrated | Preference-tuning residue in the base model | Prompting or a different model | See Preference Tuning |
The strongest honest case for fine-tuning is the second-to-last row plus Model Selection and Routing: a small tuned model on a narrow task can beat a large prompted one on cost and latency and sometimes accuracy, because the task is narrow enough that generality buys nothing.
Full Fine-Tuning vs Parameter-Efficient Methods
Full fine-tuning updates every weight. You get maximum capacity to change behaviour and you pay for it: memory for the weights plus optimizer state, a complete model copy per variant, and a real risk of degrading capabilities you never intended to touch.
LoRA β low-rank adaptation β is the mechanism worth actually understanding. Instead of updating a large weight matrix, you freeze it and learn a small delta expressed as the product of two skinny matrices. Because the delta is constrained to be low-rank, it has orders of magnitude fewer trainable parameters than the layer it modifies. At inference the frozen base output and the adapter output are combined, so the model behaves as if the weights changed while the base file on disk never did.
flowchart LR
I[Input tokens] --> B[Frozen base weights - never updated]
I --> L[LoRA adapter - two small low rank matrices]
B --> M[Combine base output with adapter delta]
L --> M
M --> O[Output tokens]
T[Training pass - gradients touch the adapter only] --> L
A[Adapter store - one small file per task or tenant] --> L
classDef client fill:#f97316,stroke:#c2410c,color:#fff
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef async fill:#b4f,stroke:#333,color:#000
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class I client
class B,M service
class L async
class T,A data
class O client
QLoRA is LoRA on top of a quantized base: hold the frozen weights at lower precision to cut memory, train the adapter in higher precision. This is what makes tuning a large open-weight model feasible on modest hardware. The precision tradeoff is the same one discussed in Distillation and Small Models.
The practical consequences matter more than the math:
- Cheap to train. Far fewer trainable parameters means less memory and shorter runs, so experiments cost hours instead of weeks.
- Cheap to store. An adapter is a small file next to one shared base, not another full model. Adapter artifacts live happily in object storage.
- Swappable at serve time. One loaded base can serve many adapters, which is what makes per-tenant and per-task variants economically sane instead of a fleet of model copies. Serving mechanics are in Inference Serving.
- Reversible. Detaching an adapter restores base behaviour exactly. Rolling back a full fine-tune means redeploying a different multi-gigabyte artifact.
| Approach | Trainable parameters | Storage per variant | Best for | Main risk |
|---|---|---|---|---|
| Full fine-tuning | All of them | A whole model | Large behaviour shifts, a new domain register, when you own the serving stack anyway | Cost, catastrophic forgetting, slow iteration |
| LoRA | A small fraction | A small adapter file | Almost every product use case | Less capacity for very large behaviour changes |
| QLoRA | Same as LoRA | Same as LoRA | Tuning bigger bases on limited memory | Quantization of the base costs some quality |
| Provider-hosted tuning | Providerβs choice | Providerβs artifact | Teams with no GPU story | Opaque internals, lock-in, limited control |
Start with LoRA. Reach for full fine-tuning only when a LoRA run measurably plateaus below your bar.
Dataset Construction Is The Real Work
Everything above is a weekend of reading. The dataset is the project. Budget accordingly: most of the calendar time on a successful tuning effort goes into examples, not training runs.
Quality beats quantity, decisively. A few hundred to a few thousand carefully curated examples routinely outperforms a far larger scraped or generated set. The reason is mechanical: training optimises toward whatever is in the file, so a noisy set teaches noise with the same diligence it would have taught signal. Ten inconsistent examples do more damage than a hundred clean ones do good.
Your dataset encodes your failure modes. This is the sentence to remember. If half your examples hedge, you are training a hedger. If your annotators disagreed about whether to abstain, you are training incoherent abstention. If every example happens to be short, long inputs will fall apart. Nothing in the process warns you β the loss curve looks identical either way.
Format consistency is non-negotiable. Same field order, same delimiters, same system prompt, same way of rendering tool calls. Where tuning shines is format adherence, and it learns the format from the shape of your examples, including the accidents in that shape.
Curate for coverage of the hard cases. Draw from real traffic the way you build a golden set in Evals: ambiguous inputs, inputs where the right answer is a refusal, long inputs, multi-part requests, near-miss entities. A dataset drawn uniformly from happy-path traffic teaches happy-path behaviour.
Split before you train, and hold something back.
- Train β what the model sees.
- Validation β watched during training to decide when to stop.
- Test β never used for any decision during development. This is the only split that can honestly tell you whether tuning helped.
Watch for catastrophic forgetting. Train hard on one narrow distribution and general capability erodes β the model gets better at your task and worse at everything adjacent, including things you depend on. Defences: keep runs short, prefer adapters over full tuning, mix in a slice of general-purpose examples, and always keep a few off-task eval cases in your suite as a canary. If your on-task score rises while the off-task canary falls, you are trading away capability you will miss later.
Overfitting signs that should stop a run: validation loss rising while training loss falls; outputs that reproduce training examples nearly verbatim; sharp degradation on inputs phrased slightly differently from the training set; and confident behaviour on inputs the model has no business being confident about.
Evaluation - The Rule You Cannot Skip
You must compare the tuned model against a prompted baseline on the same eval set, or you cannot claim the tuning helped. This is the single most-violated rule in applied fine-tuning, and skipping it is how teams end up maintaining a model artifact that a better prompt would have matched.
A defensible comparison holds everything constant except the thing you changed:
- Freeze the eval set before training starts. Building it afterwards invites you to build the set the tuned model wins on.
- Baseline with real effort. The baseline is your best prompt, not the lazy one you used before you got excited about tuning. Include the few-shot examples.
- Score on the held-out test split, using the same grading path β assertions, reference metrics, and a validated judge if you have one.
- Compare on all four axes: quality, latency, cost per request, and operational burden. Tuning that wins on quality and loses on the other three has not obviously won.
- Confirm the movement exceeds run-to-run noise. Sampled generation is stochastic; one lucky run is not a result.
- Check the off-task canary to catch capability you quietly gave up.
If the tuned model wins by a small margin, the prompted baseline is usually still the better engineering choice, because it has no artifact to own.
Serving and Lifecycle - You Now Own a Model
The moment tuning succeeds, you have acquired an asset with ongoing obligations. This is the cost nobody puts in the proposal.
- Versioning. Every artifact needs a version, the dataset snapshot that produced it, the base model it was built on, and the eval scores it shipped with. An artifact you cannot reproduce is an artifact you cannot debug.
- Base model churn. Your tuning is relative to a base. When that base is deprecated or upgraded, your work does not transfer β you re-tune and re-evaluate.
- Hosting. Provider-hosted tuning bills differently from base inference. Self-hosting means capacity planning, which is GPU Capacity and Cost Engineering territory.
- Rollout and rollback. Treat a model swap as a deployment: flag it, canary it on a traffic slice, define the rollback trigger up front, and keep the previous artifact warm. Standard deployment reliability practice, with an eval score as the canary metric.
- Drift. The data distribution moves. A tuned model is a photograph of the task as it looked when you curated the set.
Bad to Good to Great
Bad - fine-tune to fix a knowledge gap
The model does not know your internal docs, so you tune on a dump of those docs. Some facts land, many do not, nothing is citable, you cannot tell which answers are grounded, and the day a policy changes you are back to the training pipeline. You have spent weeks building an expensive, un-auditable, already-stale search index inside a weight matrix.
Good - LoRA on a curated set, shipped without a baseline
You classify the failure correctly as a behaviour problem, curate a few hundred clean examples, train an adapter, and the outputs look much better. This is real progress. The gap is that βmuch betterβ is measured against the old prompt, so you do not know whether the win came from tuning or from finally thinking clearly about the task. You are now maintaining an artifact whose value is unquantified.
Great - classified, prompted-baseline-beaten, versioned, and monitored
- Failure classified from error analysis β tuning chosen because the symptom is behaviour, format, tone, or unit economics, not missing facts.
- Prompting and retrieval exhausted first, and the remaining gap written down as a number.
- Dataset curated from real traffic, format-consistent, hard cases over-sampled, with train, validation, and test splits fixed before training.
- LoRA or QLoRA first, full fine-tuning only after an adapter demonstrably plateaus.
- Head-to-head against the best prompted baseline on a frozen test split across quality, latency, cost, and operational burden.
- Off-task canary cases in the suite to catch forgetting.
- Artifact versioned with its dataset snapshot, base model, and eval scores, rolled out behind a flag with a defined rollback.
- A re-tuning trigger defined in advance β base deprecation, drift, or an eval regression β so the model has a maintenance plan rather than an owner who left.
When to Use
β Fine-tune when:
- The behaviour you want is a habit, format, or tone that prompting reaches inconsistently
- A narrow high-volume task would run acceptably on a small tuned model at much lower cost and latency
- You are paying for a long few-shot block on every single request and want to absorb it into weights
- The domain has conventions and shorthand that cost hundreds of prompt tokens to explain every call
- You need per-tenant or per-task behaviour variants and adapters make that cheap to serve
β Do not fine-tune when:
- The problem is missing, private, or changing facts β that is retrieval, and it wins on cost, freshness, and citability
- You have not yet built an eval set, because you will have no way to know whether it worked
- Your best prompt is genuinely untested and your baseline is a first draft
- The output format is the only issue and schema-constrained decoding has not been tried β see Structured Outputs
- You have fewer than a few dozen clean examples and no plan to collect more
- Nobody is willing to own the artifact, its versioning, and its eventual re-tune
Common Interview Questions
Q1: When would you fine-tune instead of using RAG?
They solve different problems, so the question is which failure I am looking at. Missing, private, or changing facts go to retrieval every time β it is cheaper, updates by re-indexing, and produces citations a tuned model structurally cannot. I fine-tune when the gap is behavioural: a format the model drifts off, a house tone that will not stick from a description, domain conventions that cost hundreds of tokens to restate each call, or a narrow high-volume task where a small tuned model beats a large prompted one on cost and latency. The two also compose β tuning the model to follow a grounding and citation discipline, while retrieval supplies the actual evidence.
Q2: Explain LoRA and why teams prefer it to full fine-tuning.
Full fine-tuning updates every weight, which needs memory for weights plus optimizer state and produces a complete model copy per variant. LoRA freezes the base and instead learns a small low-rank delta β two skinny matrices whose product stands in for the weight update β so the number of trainable parameters drops by orders of magnitude. At inference the base output and the adapter output are combined, so behaviour changes while the base file does not. The practical wins are what matter: runs are short enough to iterate, each variant is a small file rather than a whole model, one loaded base can serve many adapters so per-tenant variants become affordable, and detaching an adapter is an exact rollback. QLoRA extends this by keeping the frozen base at lower precision to fit bigger models in less memory.
Q3: How many examples do you need and how do you build them?
Fewer than people expect, and cleaner than people manage β a few hundred to a few thousand well-curated examples routinely beats a much larger noisy set, because training optimises toward whatever is in the file. I source from real traffic rather than imagination, since invented examples are systematically too polite and too on-topic. I enforce format consistency ruthlessly, because format is exactly what the model learns best and it learns the accidents too. I over-sample the hard cases β ambiguous inputs, inputs where refusal is correct, long inputs β then fix train, validation, and test splits before training and never let a development decision touch the test split. The framing I keep in mind is that the dataset encodes my failure modes: if my examples hedge, I am training a hedger, and nothing in the loss curve will tell me.
Q4: You fine-tuned and the eval score went up. What would you check before shipping?
Whether the comparison was fair. The tuned model has to beat my best prompted baseline on the same frozen eval set, not the weak prompt I started with, or I cannot attribute the gain to tuning at all. Then: did the win hold on the held-out test split rather than the validation split I was steering against; is the movement bigger than run-to-run sampling noise; did my off-task canary cases stay flat, since a rise on-task with a drop off-task means catastrophic forgetting I will regret; and did latency, cost per request, and operational burden stay inside budget. A small quality win that adds a model artifact to own is usually a net loss.
Q5: What are the hidden costs of a fine-tuned model in production?
Owning an artifact. It needs a version bound to the dataset snapshot, the base model, and the eval scores it shipped with, or it is undebuggable. It is anchored to a base model that will eventually be deprecated or upgraded, and the tuning does not transfer β you re-tune and re-evaluate. Serving is either a provider bill at different rates or your own capacity planning. A model swap is a deployment, so it needs a flag, a canary, and a rollback with the old artifact kept warm. And the task distribution drifts, so the artifact silently ages. None of that exists with a prompt you can edit and redeploy in a minute, which is why the prompted baseline deserves a genuine chance to win.
Build it in code: Agentic AI Course · Fundamentals: Core Concepts