RLHF, DPO, and Preference Tuning - Complete Deep Dive
Prerequisites: Fine-Tuning, How LLMs Actually Work, Evals Used in: LLM-as-Judge, Model Selection and Routing, Distillation and Small Models
What is Preference Tuning?
Preference tuning is training a model on judgements about which of two outputs is better, rather than on a single demonstrated correct answer. You do not tell the model what to say. You show it pairs β this response was preferred over that one β and the training process shifts the model toward the kind of output people picked.
It sits one layer past ordinary fine-tuning, and it is even further down the list of things a product engineer should reach for. Prompting first, retrieval for knowledge, evals to know whether anything helped, supervised fine-tuning for behaviour and format. Preference tuning is for the case where you cannot write down the right answer but you can reliably recognise the better of two.
Be honest about who this page is for. Most AI engineers consume preference-tuned models rather than producing them. Every major hosted model you call has already been through this process, tuned to somebody elseβs preferences by somebody elseβs raters. The practical value here is diagnostic: when a model pads its answers, agrees with you when you are wrong, refuses something harmless, or hedges a clear-cut fact, you are looking at preference-tuning residue. Recognising that tells you to fix it with prompting or model choice, not with a training run of your own.
Real-world analogy: Supervised fine-tuning is handing a new writer one model memo and saying βwrite like this.β Preference tuning is what an editor actually does β reads two drafts and says βthis one, and hereβs roughly why,β a few thousand times, until the writer internalises the house taste. No single memo could have captured that taste, because taste is comparative.
Why Supervised Fine-Tuning Is Not Enough
Supervised fine-tuning teaches imitation of one demonstrated answer. That works when a right answer exists. For a large share of real tasks it does not.
Ask a model to reply to an annoyed customer, summarise a messy thread, explain a tradeoff, or decide how much caution an ambiguous request deserves. There is no single correct output. There are better and worse ones, and the difference lives in things a single demonstration cannot express: how much to say, when to ask instead of answer, how confident to sound, when to decline.
Three specific limits:
- The one-answer assumption breaks. SFT pushes probability toward the exact text in your file, so of the many good answers, only the phrasing your annotator happened to write is reinforced. Equally good alternatives get pushed down.
- You can label preference far more cheaply than you can write gold answers. Asking a rater to write an excellent response to a hard support ticket is slow and needs an expert. Asking them which of two responses is better is fast and needs only judgement. The cheaper signal is also the more available one.
- SFT cannot express what to avoid. A demonstration says βdo this.β It has no way to say βand specifically not thatβ β no negative signal at all. Preference pairs carry both.
So the data you can realistically collect is not a set of ideal answers. It is a pile of comparisons. Preference tuning is the family of methods that learns from that shape of data directly.
The RLHF Pipeline, Conceptually
Reinforcement learning from human feedback is three stages stacked on a model that has already been supervised fine-tuned.
- Collect pairwise preferences. Sample two or more responses to the same prompt, have humans pick the better one. The output is a dataset of prompt plus preferred plus rejected.
- Train a reward model. A separate model learns to predict which response humans would prefer, turning scattered human judgements into a scoring function you can call millions of times without paying a rater.
- Optimize the policy against that reward, with a constraint. The model being tuned β the policy β generates, gets scored by the reward model, and updates to earn higher scores. Crucially, a penalty holds it near the original model it started from.
The constraint is the part people skip when explaining this, and it is what makes the method work at all. The reward model is a proxy for human preference, accurate on the kind of output it was trained on and unreliable elsewhere. Unconstrained optimization against a proxy does not find great answers β it finds the proxyβs blind spots. Let it run free and the policy drifts into degenerate territory: repetitive phrasings that happen to score well, collapsing diversity, fluency decaying while the reward number climbs. Anchoring the policy to its starting point keeps it inside the region where the reward modelβs opinion still means something. The tuning strength becomes a dial between βbarely changedβ and βhigh reward and visibly broken.β
Reward Hacking - The Central Failure Mode
Everything that goes wrong here is a version of one thing: the policy learns to satisfy the measurement instead of the intent. The reward model was supposed to stand in for βa person would find this helpful.β What it actually encodes is βoutput that resembles what raters clicked on.β Optimize hard enough and the gap becomes the strategy.
It shows up as recognisable, familiar behaviour:
| Observed behaviour | What the proxy rewarded | Why raters produced it |
|---|---|---|
| Verbosity and padding - restating the question, bulleting the obvious, closing summaries | Length correlates with effort | Longer answers look more thorough at a glance |
| Sycophancy - folding when a user pushes back, even on a correct answer | Agreement correlates with satisfaction | Raters prefer responses that validate them |
| Hedging - qualifying claims that do not need qualifying | Caution rarely gets marked wrong | Safe non-answers avoid penalties |
| Confident tone regardless of grounding - authoritative delivery of shaky content | Fluency and assurance read as competence | Confidence is easier to judge than correctness |
| Over-refusal - declining harmless requests with surface-level resemblance to unsafe ones | Refusal is never penalised as harm | Raters penalise unsafe output far harder than unhelpful output |
| Formatting theatre - headings and structure on a one-line answer | Structure looks organised | Formatted text scans as higher quality |
This is Goodhartβs law with a training loop. Note that none of these are bugs in the code β the optimization worked exactly as specified, against a specification that was subtly wrong. The two defences are the constraint keeping the policy near its origin, and never trusting the reward number alone: you evaluate against held-out human judgement, because a rising reward score with flat human preference is the signature of hacking. The same trap applies to automated graders, which is why LLM-as-Judge insists a judge be validated against human labels rather than believed.
DPO - The Pragmatic Shortcut
Direct preference optimization asks: if the reward model is only ever used to rank outputs, and the ranking information is already sitting in the preference pairs, why train the intermediate model at all? DPO optimizes the policy directly on the preference pairs β raise the likelihood of the preferred response, lower the likelihood of the rejected one, while a reference copy of the starting model keeps the policy anchored the same way the RLHF constraint does.
flowchart LR
subgraph RLHF["RLHF - three moving parts"]
A[Pairwise preference data] --> B[Train a reward model]
B --> C[Optimize the policy against the reward]
R1[Frozen reference model] --> C
C --> D[Tuned policy]
end
subgraph DPO["DPO - one moving part"]
E[Same pairwise preference data] --> F[Optimize the policy on the pairs directly]
R2[Frozen reference model] --> F
F --> G[Tuned policy]
end
classDef client fill:#f97316,stroke:#c2410c,color:#fff
classDef service fill:#10b981,stroke:#065f46,color:#fff
classDef async fill:#b4f,stroke:#333,color:#000
classDef data fill:#fbbf24,stroke:#92400e,color:#000
class A,E data
class B async
class C,F service
class R1,R2 data
class D,G client
What this buys, in engineering terms:
- One model to train instead of two. No separate reward model to build, validate, version, or keep in sync.
- No RL loop. Training looks like ordinary supervised training over pairs, which means it is far more stable and far easier to debug. Anyone who has run an RL loop appreciates what removing it is worth.
- Fewer knobs, fewer ways to fail. The most common RLHF failures are reward-model and optimization-loop failures. Deleting both components deletes those failure classes.
What you give up:
- Flexibility. A reward model is a reusable scoring function β you can score arbitrary new outputs, rank candidates at inference, or filter a generated dataset with it. DPO produces a tuned policy and nothing else.
- Headroom on the same data. A well-built RLHF setup can keep generating fresh samples and scoring them, extracting more than a fixed pair set contains. DPO learns from the pairs you have.
- The proxy problem does not disappear. DPO removes the reward model, not the fact that preference data is a proxy for intent. Sycophancy and verbosity are perfectly learnable from pairs.
A family of related direct-preference methods exists, differing in how the objective is framed and whether a reference model is needed at all. The distinctions matter to people training models; for product work, knowing that the category exists is enough.
SFT vs RLHF vs DPO
| Β | Supervised fine-tuning | RLHF | DPO |
|---|---|---|---|
| Data it needs | Prompt plus one good answer | Pairwise preferences over sampled responses | Pairwise preferences over sampled responses |
| What it buys | Imitation of a demonstrated style, format, or task | Nuanced quality judgement plus a reusable scoring function | Most of RLHFβs quality gain at a fraction of the complexity |
| Moving parts | One training run | Reward model plus RL loop plus reference model | One training run plus a reference model |
| Operational complexity | Low | High - hardest of the three to run and debug | Moderate |
| Signature failure | Learns the accidents of your examples and only the phrasing you wrote | Reward hacking - verbosity sycophancy hedging over-refusal | Inherits rater bias and can overfit the pair set |
| Reasonable default for | Format tone and narrow tasks | Labs with rater infrastructure and a research budget | Teams that genuinely need preference tuning and want it to finish |
Preference Data Is The Expensive Part
The methods are published. The data is the moat, and the cost.
Rater agreement is the ceiling on everything. If two qualified raters disagree about which response is better, no training method can learn a coherent preference from their combined labels β it learns their average, which may be a behaviour neither would endorse. Measure inter-rater agreement before collecting at volume. Low agreement is not a data problem to be averaged away, it is a sign the guideline does not say what you meant.
Your ratersβ biases become the modelβs values. This is the sentence worth carrying out of this page. Raters under time pressure prefer longer answers, confident tone, and validation β every documented reward-hacking behaviour traces back to something a human actually clicked. If your guideline does not explicitly reward concision, you are shipping verbosity. If it does not distinguish βcorrect and unwelcomeβ from βwrong,β you are shipping sycophancy. The guideline is the product specification for the modelβs personality.
The practical mechanics that decide whether the data is usable:
- Write the guideline first, pilot it on a small batch, and rewrite it once raters have argued about real cases.
- Define tiebreakers explicitly β what wins when one response is more accurate and the other more helpful.
- Sample the response pairs from the model you will actually tune, on prompts from real traffic. Pairs drawn from a different distribution teach preferences about a system you are not shipping.
- Over-sample the contested middle. Pairs where one response is obviously terrible teach almost nothing.
- Keep a held-out human-labelled slice you never train on. It is the only instrument that can detect reward hacking.
Diagnosing Preference-Tuning Residue
Here is the part that pays off for a product engineer. The base model you are calling has already been preference-tuned, so some of what you read as a prompting failure is actually inherited alignment behaviour β and knowing that changes the fix.
| What you observe | Likely residue | What to do about it |
|---|---|---|
| Answers are padded and restate your question | Length-correlated reward | Ask explicitly for brevity, constrain output length, set a format budget in the prompt |
| Model reverses a correct answer when you push back | Sycophancy | Instruct it to hold positions under challenge and cite evidence, or route to a model that holds firmer |
| Harmless requests get refused | Over-refusal from asymmetric safety penalties | Rephrase intent and context, supply a legitimate-use framing, or change model - see Model Selection and Routing |
| Clear facts arrive wrapped in qualifiers | Hedging reward | Ask for a direct answer plus a separate confidence note so caution has somewhere to live |
| Confident phrasing on ungrounded claims | Fluency read as competence | Force grounding and citation, and evaluate faithfulness rather than reading tone as truth |
| Every answer gets headings and bullets | Formatting theatre | Specify the output shape, or constrain it with a schema |
The general rule: behaviour that appears on every prompt, across your whole product, and resists prompt edits is more likely a property of the model than a flaw in your prompt. At that point the lever is model choice, not another prompt revision. Which of these are real and which are your own imagination is a question only an eval set answers β see Evals.
Bad to Good to Great
Bad - blame the prompt forever
The model hedges and pads, so you add βbe concise and directβ and iterate for a week. You are fighting an optimization that ran over far more data than your prompt contains. Some of it yields to prompting; the part that does not will not, and you have no way to tell which is which because nothing is measured.
Good - recognise the residue and work around it
You classify the behaviour as inherited, prompt against the parts that respond, and constrain what you can with output schemas and length budgets. Sensible, and enough for most products. The gap is that βseems better nowβ is still a vibe β you have not measured the behaviour, so you will not notice when a model update changes it.
Great - measure the behaviour and treat it as a model-selection criterion
- Turn each suspected residue into an eval case β a pushback case for sycophancy, a benign-but-edgy case for over-refusal, a clear-fact case for hedging, a length check for padding.
- Score candidate models on that suite, so βthis model is less sycophanticβ becomes a number you can defend.
- Prompt and constrain the part that responds to prompting, and re-measure rather than assuming.
- Make it a routing decision where models differ meaningfully on cases you care about.
- Keep the suite running, because a provider updating a model behind a stable name can change all of this without telling you.
- Only then consider tuning your own preferences β with rater guidelines written first, agreement measured, DPO before RLHF, and a held-out human-labelled slice to catch the policy gaming your proxy.
When to Use
β Worth understanding or doing when:
- You are diagnosing why a model pads, hedges, flatters, or over-refuses and need to know whether prompting can fix it
- You are choosing between models on behavioural grounds and want criteria beyond raw capability
- Your task genuinely has no single right answer but reliable comparative judgement is available
- You already have preference data as a byproduct β accepted versus rejected drafts, chosen suggestions, edited versus shipped text
- You are building an automated grader and need to understand why an unvalidated proxy drifts
β Do not run your own preference tuning when:
- Supervised fine-tuning has not been tried and the behaviour is really about format or tone
- The gap is knowledge, not judgement β that is retrieval, and it is not close
- You cannot get consistent raters, since low agreement guarantees incoherent preferences
- You have no held-out human-labelled slice, which means no way to detect reward hacking
- You are chasing a behaviour that a different base model already exhibits, which is a model-selection decision, not a training project
Common Interview Questions
Q1: Why is supervised fine-tuning not enough for quality?
SFT teaches imitation of one demonstrated answer, which assumes a single right answer exists. For tone, caution, length, when to ask instead of answer, and how confident to sound, it does not β there are better and worse responses, and the difference is comparative. SFT also has no negative signal: a demonstration can say βdo thisβ but never βspecifically not that.β And it fights the data you can actually collect, since asking a rater to pick the better of two responses is far cheaper and more reliable than asking them to author an ideal one. Preference methods learn from comparisons, which is the shape the available signal actually has.
Q2: Walk through RLHF and explain why the constraint on the policy matters.
Three stages on top of an SFT model. Collect pairwise human preferences over sampled responses. Train a reward model to predict which response humans would prefer, which converts scattered judgements into a scoring function callable millions of times. Then optimize the policy against that reward while a penalty keeps it near the model it started from. The constraint is the whole ballgame, because the reward model is a proxy that is only accurate near the distribution it was trained on. Optimize without the anchor and the policy does not find better answers, it finds the proxyβs blind spots β degenerate repetitive phrasings that score well, collapsing output diversity, reward climbing while real quality falls. The anchor keeps the policy where the reward modelβs opinion is still worth something, which makes the tuning strength a dial between barely-changed and high-reward-and-broken.
Q3: What is reward hacking and how does it show up in a shipped model?
The policy learns to satisfy the measurement rather than the intent, because the reward model encodes βwhat raters clickedβ rather than βwhat actually helps.β You can read it straight off a deployed model. Verbosity and padding, because length looks like effort. Sycophancy, because agreement reads as satisfaction and raters reward validation. Hedging, because caution is rarely marked wrong. Confident tone on ungrounded content, because fluency is easier to judge than correctness. Over-refusal, because unsafe output is penalised much harder than unhelpful output. Formatting theatre, because structure scans as quality. The defences are the anchor to the reference model, and evaluating against held-out human judgement β reward score rising while human preference stays flat is the signature.
Q4: What does DPO change relative to RLHF, and what does it give up?
DPO skips the reward model and the RL loop, optimizing the policy directly on preference pairs β raise the preferred responseβs likelihood, lower the rejected oneβs, with a frozen reference model providing the same anchoring role the RLHF penalty played. That deletes an entire model you would otherwise have to build, validate, and version, and replaces an unstable RL loop with something that trains like ordinary supervised learning. The cost is flexibility. A reward model is a reusable scorer you can point at any new output to rank candidates or filter a dataset; DPO leaves you with a tuned policy and nothing reusable. RLHF can also keep sampling and scoring fresh outputs, so it has more headroom than a fixed pair set. And DPO removes the reward model, not the proxy problem β verbosity and sycophancy are entirely learnable from pairs.
Q5: As a product engineer who will never train a reward model, what is this worth to you?
Diagnosis. Every hosted model I call has been preference-tuned to someone elseβs rater guidelines, so a meaningful share of what looks like a prompting failure is inherited behaviour. Padding, hedging, folding under pushback, and refusing harmless requests are recognisable residue with known causes, and knowing the cause changes the fix: the lever is prompting and output constraints for the part that responds, and model choice for the part that does not. The tell is behaviour that appears across every prompt in the product and survives prompt edits β at that point I stop revising the prompt and start comparing models. Concretely, I turn each suspected residue into an eval case so βless sycophanticβ becomes a number, and keep that suite running, because a provider can change all of it behind an unchanged model name.
Build it in code: Agentic AI Course · Fundamentals: Core Concepts