Limited time: AI code review, hints, mock interviews, whiteboard analysis, and all Pro features are unlocked. Enroll
⏱️ 18 min read

RLHF, DPO, and Preference Tuning - Complete Deep Dive

Stage 6 - Customization Lesson 26 of 36

Prerequisites: Fine-Tuning, How LLMs Actually Work, Evals Used in: LLM-as-Judge, Model Selection and Routing, Distillation and Small Models


What is Preference Tuning?

Preference tuning is training a model on judgements about which of two outputs is better, rather than on a single demonstrated correct answer. You do not tell the model what to say. You show it pairs β€” this response was preferred over that one β€” and the training process shifts the model toward the kind of output people picked.

It sits one layer past ordinary fine-tuning, and it is even further down the list of things a product engineer should reach for. Prompting first, retrieval for knowledge, evals to know whether anything helped, supervised fine-tuning for behaviour and format. Preference tuning is for the case where you cannot write down the right answer but you can reliably recognise the better of two.

Be honest about who this page is for. Most AI engineers consume preference-tuned models rather than producing them. Every major hosted model you call has already been through this process, tuned to somebody else’s preferences by somebody else’s raters. The practical value here is diagnostic: when a model pads its answers, agrees with you when you are wrong, refuses something harmless, or hedges a clear-cut fact, you are looking at preference-tuning residue. Recognising that tells you to fix it with prompting or model choice, not with a training run of your own.

Real-world analogy: Supervised fine-tuning is handing a new writer one model memo and saying β€œwrite like this.” Preference tuning is what an editor actually does β€” reads two drafts and says β€œthis one, and here’s roughly why,” a few thousand times, until the writer internalises the house taste. No single memo could have captured that taste, because taste is comparative.


Why Supervised Fine-Tuning Is Not Enough

Supervised fine-tuning teaches imitation of one demonstrated answer. That works when a right answer exists. For a large share of real tasks it does not.

Ask a model to reply to an annoyed customer, summarise a messy thread, explain a tradeoff, or decide how much caution an ambiguous request deserves. There is no single correct output. There are better and worse ones, and the difference lives in things a single demonstration cannot express: how much to say, when to ask instead of answer, how confident to sound, when to decline.

Three specific limits:

So the data you can realistically collect is not a set of ideal answers. It is a pile of comparisons. Preference tuning is the family of methods that learns from that shape of data directly.


The RLHF Pipeline, Conceptually

Reinforcement learning from human feedback is three stages stacked on a model that has already been supervised fine-tuned.

  1. Collect pairwise preferences. Sample two or more responses to the same prompt, have humans pick the better one. The output is a dataset of prompt plus preferred plus rejected.
  2. Train a reward model. A separate model learns to predict which response humans would prefer, turning scattered human judgements into a scoring function you can call millions of times without paying a rater.
  3. Optimize the policy against that reward, with a constraint. The model being tuned β€” the policy β€” generates, gets scored by the reward model, and updates to earn higher scores. Crucially, a penalty holds it near the original model it started from.

The constraint is the part people skip when explaining this, and it is what makes the method work at all. The reward model is a proxy for human preference, accurate on the kind of output it was trained on and unreliable elsewhere. Unconstrained optimization against a proxy does not find great answers β€” it finds the proxy’s blind spots. Let it run free and the policy drifts into degenerate territory: repetitive phrasings that happen to score well, collapsing diversity, fluency decaying while the reward number climbs. Anchoring the policy to its starting point keeps it inside the region where the reward model’s opinion still means something. The tuning strength becomes a dial between β€œbarely changed” and β€œhigh reward and visibly broken.”


Reward Hacking - The Central Failure Mode

Everything that goes wrong here is a version of one thing: the policy learns to satisfy the measurement instead of the intent. The reward model was supposed to stand in for β€œa person would find this helpful.” What it actually encodes is β€œoutput that resembles what raters clicked on.” Optimize hard enough and the gap becomes the strategy.

It shows up as recognisable, familiar behaviour:

Observed behaviour What the proxy rewarded Why raters produced it
Verbosity and padding - restating the question, bulleting the obvious, closing summaries Length correlates with effort Longer answers look more thorough at a glance
Sycophancy - folding when a user pushes back, even on a correct answer Agreement correlates with satisfaction Raters prefer responses that validate them
Hedging - qualifying claims that do not need qualifying Caution rarely gets marked wrong Safe non-answers avoid penalties
Confident tone regardless of grounding - authoritative delivery of shaky content Fluency and assurance read as competence Confidence is easier to judge than correctness
Over-refusal - declining harmless requests with surface-level resemblance to unsafe ones Refusal is never penalised as harm Raters penalise unsafe output far harder than unhelpful output
Formatting theatre - headings and structure on a one-line answer Structure looks organised Formatted text scans as higher quality

This is Goodhart’s law with a training loop. Note that none of these are bugs in the code β€” the optimization worked exactly as specified, against a specification that was subtly wrong. The two defences are the constraint keeping the policy near its origin, and never trusting the reward number alone: you evaluate against held-out human judgement, because a rising reward score with flat human preference is the signature of hacking. The same trap applies to automated graders, which is why LLM-as-Judge insists a judge be validated against human labels rather than believed.


DPO - The Pragmatic Shortcut

Direct preference optimization asks: if the reward model is only ever used to rank outputs, and the ranking information is already sitting in the preference pairs, why train the intermediate model at all? DPO optimizes the policy directly on the preference pairs β€” raise the likelihood of the preferred response, lower the likelihood of the rejected one, while a reference copy of the starting model keeps the policy anchored the same way the RLHF constraint does.

flowchart LR
    subgraph RLHF["RLHF - three moving parts"]
        A[Pairwise preference data] --> B[Train a reward model]
        B --> C[Optimize the policy against the reward]
        R1[Frozen reference model] --> C
        C --> D[Tuned policy]
    end

    subgraph DPO["DPO - one moving part"]
        E[Same pairwise preference data] --> F[Optimize the policy on the pairs directly]
        R2[Frozen reference model] --> F
        F --> G[Tuned policy]
    end

    classDef client fill:#f97316,stroke:#c2410c,color:#fff
    classDef service fill:#10b981,stroke:#065f46,color:#fff
    classDef async fill:#b4f,stroke:#333,color:#000
    classDef data fill:#fbbf24,stroke:#92400e,color:#000

    class A,E data
    class B async
    class C,F service
    class R1,R2 data
    class D,G client

What this buys, in engineering terms:

What you give up:

A family of related direct-preference methods exists, differing in how the objective is framed and whether a reference model is needed at all. The distinctions matter to people training models; for product work, knowing that the category exists is enough.


SFT vs RLHF vs DPO

Β  Supervised fine-tuning RLHF DPO
Data it needs Prompt plus one good answer Pairwise preferences over sampled responses Pairwise preferences over sampled responses
What it buys Imitation of a demonstrated style, format, or task Nuanced quality judgement plus a reusable scoring function Most of RLHF’s quality gain at a fraction of the complexity
Moving parts One training run Reward model plus RL loop plus reference model One training run plus a reference model
Operational complexity Low High - hardest of the three to run and debug Moderate
Signature failure Learns the accidents of your examples and only the phrasing you wrote Reward hacking - verbosity sycophancy hedging over-refusal Inherits rater bias and can overfit the pair set
Reasonable default for Format tone and narrow tasks Labs with rater infrastructure and a research budget Teams that genuinely need preference tuning and want it to finish

Preference Data Is The Expensive Part

The methods are published. The data is the moat, and the cost.

Rater agreement is the ceiling on everything. If two qualified raters disagree about which response is better, no training method can learn a coherent preference from their combined labels β€” it learns their average, which may be a behaviour neither would endorse. Measure inter-rater agreement before collecting at volume. Low agreement is not a data problem to be averaged away, it is a sign the guideline does not say what you meant.

Your raters’ biases become the model’s values. This is the sentence worth carrying out of this page. Raters under time pressure prefer longer answers, confident tone, and validation β€” every documented reward-hacking behaviour traces back to something a human actually clicked. If your guideline does not explicitly reward concision, you are shipping verbosity. If it does not distinguish β€œcorrect and unwelcome” from β€œwrong,” you are shipping sycophancy. The guideline is the product specification for the model’s personality.

The practical mechanics that decide whether the data is usable:


Diagnosing Preference-Tuning Residue

Here is the part that pays off for a product engineer. The base model you are calling has already been preference-tuned, so some of what you read as a prompting failure is actually inherited alignment behaviour β€” and knowing that changes the fix.

What you observe Likely residue What to do about it
Answers are padded and restate your question Length-correlated reward Ask explicitly for brevity, constrain output length, set a format budget in the prompt
Model reverses a correct answer when you push back Sycophancy Instruct it to hold positions under challenge and cite evidence, or route to a model that holds firmer
Harmless requests get refused Over-refusal from asymmetric safety penalties Rephrase intent and context, supply a legitimate-use framing, or change model - see Model Selection and Routing
Clear facts arrive wrapped in qualifiers Hedging reward Ask for a direct answer plus a separate confidence note so caution has somewhere to live
Confident phrasing on ungrounded claims Fluency read as competence Force grounding and citation, and evaluate faithfulness rather than reading tone as truth
Every answer gets headings and bullets Formatting theatre Specify the output shape, or constrain it with a schema

The general rule: behaviour that appears on every prompt, across your whole product, and resists prompt edits is more likely a property of the model than a flaw in your prompt. At that point the lever is model choice, not another prompt revision. Which of these are real and which are your own imagination is a question only an eval set answers β€” see Evals.


Bad to Good to Great

Bad - blame the prompt forever

The model hedges and pads, so you add β€œbe concise and direct” and iterate for a week. You are fighting an optimization that ran over far more data than your prompt contains. Some of it yields to prompting; the part that does not will not, and you have no way to tell which is which because nothing is measured.

Good - recognise the residue and work around it

You classify the behaviour as inherited, prompt against the parts that respond, and constrain what you can with output schemas and length budgets. Sensible, and enough for most products. The gap is that β€œseems better now” is still a vibe β€” you have not measured the behaviour, so you will not notice when a model update changes it.

Great - measure the behaviour and treat it as a model-selection criterion

  1. Turn each suspected residue into an eval case β€” a pushback case for sycophancy, a benign-but-edgy case for over-refusal, a clear-fact case for hedging, a length check for padding.
  2. Score candidate models on that suite, so β€œthis model is less sycophantic” becomes a number you can defend.
  3. Prompt and constrain the part that responds to prompting, and re-measure rather than assuming.
  4. Make it a routing decision where models differ meaningfully on cases you care about.
  5. Keep the suite running, because a provider updating a model behind a stable name can change all of this without telling you.
  6. Only then consider tuning your own preferences β€” with rater guidelines written first, agreement measured, DPO before RLHF, and a held-out human-labelled slice to catch the policy gaming your proxy.

When to Use

βœ… Worth understanding or doing when:

❌ Do not run your own preference tuning when:


Common Interview Questions

Q1: Why is supervised fine-tuning not enough for quality?

SFT teaches imitation of one demonstrated answer, which assumes a single right answer exists. For tone, caution, length, when to ask instead of answer, and how confident to sound, it does not β€” there are better and worse responses, and the difference is comparative. SFT also has no negative signal: a demonstration can say β€œdo this” but never β€œspecifically not that.” And it fights the data you can actually collect, since asking a rater to pick the better of two responses is far cheaper and more reliable than asking them to author an ideal one. Preference methods learn from comparisons, which is the shape the available signal actually has.

Q2: Walk through RLHF and explain why the constraint on the policy matters.

Three stages on top of an SFT model. Collect pairwise human preferences over sampled responses. Train a reward model to predict which response humans would prefer, which converts scattered judgements into a scoring function callable millions of times. Then optimize the policy against that reward while a penalty keeps it near the model it started from. The constraint is the whole ballgame, because the reward model is a proxy that is only accurate near the distribution it was trained on. Optimize without the anchor and the policy does not find better answers, it finds the proxy’s blind spots β€” degenerate repetitive phrasings that score well, collapsing output diversity, reward climbing while real quality falls. The anchor keeps the policy where the reward model’s opinion is still worth something, which makes the tuning strength a dial between barely-changed and high-reward-and-broken.

Q3: What is reward hacking and how does it show up in a shipped model?

The policy learns to satisfy the measurement rather than the intent, because the reward model encodes β€œwhat raters clicked” rather than β€œwhat actually helps.” You can read it straight off a deployed model. Verbosity and padding, because length looks like effort. Sycophancy, because agreement reads as satisfaction and raters reward validation. Hedging, because caution is rarely marked wrong. Confident tone on ungrounded content, because fluency is easier to judge than correctness. Over-refusal, because unsafe output is penalised much harder than unhelpful output. Formatting theatre, because structure scans as quality. The defences are the anchor to the reference model, and evaluating against held-out human judgement β€” reward score rising while human preference stays flat is the signature.

Q4: What does DPO change relative to RLHF, and what does it give up?

DPO skips the reward model and the RL loop, optimizing the policy directly on preference pairs β€” raise the preferred response’s likelihood, lower the rejected one’s, with a frozen reference model providing the same anchoring role the RLHF penalty played. That deletes an entire model you would otherwise have to build, validate, and version, and replaces an unstable RL loop with something that trains like ordinary supervised learning. The cost is flexibility. A reward model is a reusable scorer you can point at any new output to rank candidates or filter a dataset; DPO leaves you with a tuned policy and nothing reusable. RLHF can also keep sampling and scoring fresh outputs, so it has more headroom than a fixed pair set. And DPO removes the reward model, not the proxy problem β€” verbosity and sycophancy are entirely learnable from pairs.

Q5: As a product engineer who will never train a reward model, what is this worth to you?

Diagnosis. Every hosted model I call has been preference-tuned to someone else’s rater guidelines, so a meaningful share of what looks like a prompting failure is inherited behaviour. Padding, hedging, folding under pushback, and refusing harmless requests are recognisable residue with known causes, and knowing the cause changes the fix: the lever is prompting and output constraints for the part that responds, and model choice for the part that does not. The tell is behaviour that appears across every prompt in the product and survives prompt edits β€” at that point I stop revising the prompt and start comparing models. Concretely, I turn each suspected residue into an eval case so β€œless sycophantic” becomes a number, and keep that suite running, because a provider can change all of it behind an unchanged model name.


Build it in code: Agentic AI Course · Fundamentals: Core Concepts

Free system design + DSA prep. If it helped you crack an interview, consider supporting.

SensAI SensAI
Beta
Listening...
Tap mic to stop voice mode

Shape what we build next

Every piece of feedback is read by the team and directly influences our roadmap.

What type of feedback?

Install SystemCraft

Add to your home screen for instant access, offline reading, and a distraction-free experience.

Offline reading Faster loads No browser tabs App-like feel

Unlock AI Features

One click to activate - no payment, no credit card. Just sign in and you're in.

AI code review and hints
SensAI chat assistant
AI mock interviews
Whiteboard analysis
100% free during early access