Lesson 4 of 6 · 48 min

Preference optimization — DPO vs RLHF

After SFT teaches format, preference tuning teaches taste — the RLHF three-stage pipeline (SFT → reward model → PPO) and its failure modes, DPO’s closed-form collapse of that pipeline into one loss, when each wins (dense preferences vs sparse shaped rewards), and how alignment regresses capabilities and diversity.

SFT teaches format; preference tuning teaches taste

SFT (L2) makes a base model follow instructions, but it can’t easily express “this answer is better than that one” — it only sees positive examples. Preference optimization is the stage that aligns the model to human preference over pairs: helpfulness, tone, refusal behaviour, the difference between an acceptable answer and a great one. The headline result — InstructGPT’s 1.3B aligned model preferred over 175B GPT-3 — comes from this stage. The modern debate is RLHF (the original three-stage pipeline) vs DPO (one supervised loss), and getting the tradeoff right is a senior alignment interview’s core.
The RLHF pipeline (Ouyang et al., 2022) is three stages, and a strong answer walks all three. (1) SFT: fine-tune the base model on curated demonstrations to get a reference policy. (2) Reward model (RM): train a scalar model r(x,y) on pairwise human preference data {y_w, y_l} (winner/loser) with a Bradley-Terry loss -log sigmoid(r(x,y_w) - r(x,y_l)) — the RM learns to score any (prompt, response). (3) PPO: optimise the policy with reinforcement learning against the RM, plus a KL penalty toward the SFT reference so the policy doesn’t drift into gibberish that games the RM. InstructGPT used ~33k prompts for the RM stage and ~31k for PPO, with a pretraining-mix term (γ ≈ 27.8) to fight regression on public NLP benchmarks.
code
1RLHF (InstructGPT) -- three stages, four models live in memory23  1. SFT          base -> reference policy   (demonstrations)4  2. REWARD MODEL r(x,y) scalar               (pairwise prefs, Bradley-Terry)5                  loss = -log sigmoid( r(x,y_w) - r(x,y_l) )6  3. PPO          maximise r(x,y) - beta·KL(policy || SFT_ref)7                  (policy + reference + reward + value models all resident)89  Failure modes: reward hacking (game the RM), PPO instability / mode collapse,10  4-model infra cost. KL penalty + pretraining-mix (gamma~27.8) fight drift.
The PPO failure modes are the part interviewers probe hardest. Reward hacking: the policy finds outputs that score high on the RM but are degenerate to humans (the RM is a proxy, and Goodhart’s law applies — “reward over-optimisation”). Instability: PPO can collapse to a narrow, high-reward, low-diversity mode. Infrastructure cost: you keep four models in memory — policy, reference, reward, and value — which is why RLHF is heavy to run. Interview angle. A weak answer stops at “RLHF aligns the model.” A strong one names reward hacking vs PPO instability vs the 4-model cost, and notes the KL penalty is the lever that keeps the policy near the SFT reference.
An update on DPO vs PPO for LLM alignmentNathan Lambert

DPO — your language model is secretly a reward model

DPO (Rafailov et al., 2023) starts from a clever observation: under the Bradley-Terry model and a KL-bounded objective, the optimal RLHF policy can be written in closed form as a function of the reference policy and the reward — which means the reward model and the policy aren’t separate objects. The reward is implicit in the log-ratio of the policy to the reference. So DPO eliminates the reward model and the PPO loop entirely, collapsing all three stages after SFT into a single supervised classification loss over (chosen, rejected) pairs. The paper’s subtitle says it: “your language model is secretly a reward model.”
code
1DPO -- one supervised loss, no reward model, no PPO sampling23  L_DPO = -log sigmoid( beta · [ log( pi(y_w|x) / pi_ref(y_w|x) )4                                 - log( pi(y_l|x) / pi_ref(y_l|x) ) ] )56  pi      = the policy being trained7  pi_ref  = the FROZEN SFT model (reference)8  beta    = KL strength (default 0.1; 0.5 for TL;DR summarization)9  y_w/y_l = chosen / rejected response1011  No reward model. No sampling from the policy during training. No PPO.12  It's a Bradley-Terry classification loss on preference pairs.
The practical consequences are why DPO took over open-weight chat alignment. There is no reward model to train, no sampling from the policy during training (PPO’s expensive, unstable rollouts are gone), and no PPO infrastructure. The paper reports DPO is “stable, performant, and computationally lightweight” and removes “the need for … significant hyperparameter tuning.” The one finicky knob is beta (the KL strength, default 0.1, 0.5 for TL;DR summarisation): lower beta lets the policy diverge more from the reference. The canonical recipe — load a 4-bit base, SFT on instructions, then DPO with beta=0.1 and a small learning rate (~1e-6) on preference pairs — is the de-facto standard for small-team and open-weight alignment.

When each wins — dense preferences vs sparse shaped rewards

The operational division mlabonne’s course states cleanly: DPO is “great for creating chat models”; PPO/RL is “ideal for creating reasoning models.” The mechanism explains why. DPO excels when the target is a dense preference over short responses — style, tone, refusal patterns — because the loss directly prefers one answer over another and the chosen/rejected signal is rich. PPO/RLHF excels when the optimisation signal is sparse and reward-shaped — math correctness, multi-step reasoning, tool-call success — because a reward model can propagate credit across a long horizon in ways a binary preference loss cannot. This is also why modern reasoning models lean on RL (and variants like GRPO) rather than pure DPO.
code
1WHICH PREFERENCE METHOD?  (match the reward signal's structure)23  Signal shape                         -> Method4  ----------------------------------------  ----------------------------5  dense preference over short answers   -> DPO (chat, tone, refusals)6  (style / helpfulness / format taste)     stable, 1 loss, no RM78  sparse, shaped, long-horizon reward   -> PPO / RLHF (reasoning, tools,9  (math correctness, multi-step, tools)     math) -- RM propagates credit1011  Calibrated reward signal needed for       -> RLHF (you get a reusable RM12  online iteration / continuous eval           for eval + online loops)1314  Most open chat models: SFT -> DPO.  Reasoning models: RL (PPO/GRPO).
The measured results bound the debate. InstructGPT: 175B InstructGPT outputs preferred over 175B GPT-3 85±3% of the time, and over few-shot 175B GPT-3 71±4%. DPO: it “exceeds PPO-based RLHF in ability to control sentiment” on IMDb, and on TL;DR summarisation hits a win rate of ~61% at temperature 0 versus PPO’s ~57% at its best sampling temperature. So DPO isn’t just cheaper — on these preference-shaped tasks it’s competitive or better. Interview angle. “When would you still reach for PPO/RLHF in 2026?” → when the signal is sparse/shaped (reasoning, tool success), when you need a calibrated reward model for online iteration and continuous eval, or when you’re running RLHF steps during deployment — not “PPO is the standard.”

How preference tuning regresses capabilities

Alignment is not free, and naming its costs is the senior move. The QDC study (L2) found that RLHF models generalise better than SFT models but exhibit worse sample diversity — alignment improves in-distribution accuracy while narrowing the output distribution. This is the “alignment tax”: a heavily preference-tuned model can become repetitive, over-hedge, or lose creative range. InstructGPT explicitly fought regression on public NLP benchmarks (SQuADv2, DROP) with a pretraining-mix term and tested KL coefficients up to 100× the default — and still reported residual regressions. The lesson: preference tuning trades expressivity for adherence, and you must measure the trade, not assume it’s positive.
DPO has its own characteristic failure, distinct in kind from PPO’s reward hacking: preference warping (and reference brittleness). Because DPO optimises against a frozen reference and a fixed set of preference pairs, it can overfit to the shape of those pairs rather than true quality — producing outputs that are stylistically aligned but factually shallow. And if the SFT reference drifts (you change it) or the pairs go stale, the DPO signal degrades; the paper itself flags out-of-distribution generalisation and reward over-optimisation in the DPO setting as open limitations. Interview angle. “How does DPO fail, and how is that different from PPO?” → PPO fails by reward hacking (gaming an explicit RM); DPO fails by preference warping and OOD brittleness against a frozen reference. Knowing the failures differ in kind is the reading-depth signal.
code
1THE ALIGNMENT TAX  (preference tuning costs something)23  Symptom                         Cause                       Mitigation4  -----------------------------   -------------------------   ----------------5  repetitive / low-diversity      RLHF narrows output dist.   sampling, rejection-6  output                          (QDC: better gen, worse D)  sampling, seed diversity7  general benchmark regression    over-optimising the reward  pretraining-mix term,8  (MMLU/SQuAD/DROP drop)          / capability drift          KL penalty, gate on eval9  stylistically aligned but       DPO preference warping      diverse, fresh prefs;10  factually shallow               (overfits pair shape)       watch OOD; don't over-train11  signal goes stale               frozen reference drifts     re-pair if you move the SFT ref1213  Always run a GENERAL eval alongside the preference eval to catch the tax.
The cross-cutting insight that ties the alignment lesson back to the data lesson: both pipelines degrade if the SFT data/reference is stale, so the lifecycle of the SFT corpus — not the choice of DPO vs PPO — is often the highest-leverage operational decision. A drifting reference quietly poisons DPO; a weak SFT stage caps everything downstream in RLHF. This is the same theme as the rest of the track: the data and the reference are the load-bearing parts, and the algorithm is the smaller decision on top.
paperDirect Preference Optimization: Your Language Model is Secretly a Reward ModelRafailov et al. (arXiv)paperTraining language models to follow instructions (InstructGPT / RLHF)Ouyang et al. (arXiv)articleFine-tune Llama 2 with DPO (the canonical SFT→DPO recipe)Hugging Face

Checkpoint

You’re aligning a customer-support chat model to a preferred tone and refusal style, with a small team and limited GPU budget. Which approach fits best and why?

AFull RLHF with PPO and a trained reward modelBSFT then DPO — a dense style/refusal preference is exactly DPO’s sweet spot: one stable loss, no reward model, no PPO infraCPrompt engineering only — never tune preferences
Sign up free to answer and see why

Checkpoint

In an interview you explain DPO. What’s the one-sentence core that signals real understanding?

ADPO doesn’t need a reward modelBUnder Bradley-Terry + a KL-bounded objective, the optimal policy is a closed-form function of the reference and reward, so the reward is implicit in the policy/reference log-ratio — letting DPO replace the RM and PPO with one classification lossCDPO uses reinforcement learning more efficiently than PPO
Sign up free to answer and see why

Checkpoint

You’re building a model that must solve multi-step math and verify tool-call success. Why might PPO/RLHF beat DPO here?

APPO trains faster than DPOBThe reward is sparse and shaped (correctness over long horizons); a reward model can propagate credit across steps in ways a binary preference loss can’tCDPO can’t be used on math problems at all
Sign up free to answer and see why

Checkpoint

After heavy DPO the model is more on-tone but its answers feel repetitive and slightly shallower factually. What’s happening and what do you do?

ANothing — preference tuning only improves the modelBIt’s the alignment tax (reduced diversity + DPO preference warping); diversify/refresh preference pairs, don’t over-train, adjust sampling, and gate on a general eval to bound the regressionCRaise beta to 1.0 to make the model more confident
Sign up free to answer and see why

Checkpoint

Which statement about RLHF failure modes vs DPO failure modes is correct?

ABoth fail identically — there’s no meaningful differenceBPPO/RLHF fails by reward hacking (gaming an explicit reward model); DPO fails by preference warping and OOD brittleness against a frozen referenceCDPO is immune to over-optimisation because it has no reward model
Sign up free to answer and see why

Interview prep

Alignment rounds test whether you can derive DPO (not just recite “no reward model”), walk the RLHF three stages with their failure modes, choose between them by the reward signal’s structure, and name the alignment tax. DPO is the quiet consensus default for chat-style alignment; the senior signal is knowing precisely when PPO/RLHF still earns its cost — and that the two methods fail in different ways.
  1. 01“Explain RLHF end-to-end.” → SFT → reward model (Bradley-Terry pairwise) → PPO with a KL penalty to the SFT reference; four models resident.
  2. 02“PPO failure modes?” → reward hacking (game the RM, Goodhart), instability/mode collapse, and the four-model infra cost.
  3. 03“How does DPO change it?” → closed-form optimal policy under BT+KL ⇒ reward implicit in the policy/reference log-ratio ⇒ one classification loss, no RM, no PPO.
  4. 04“DPO’s finicky knob?” → beta (KL strength), default 0.1 (0.5 for TL;DR); lower beta = more divergence from the reference.
  5. 05“When DPO vs PPO?” → DPO for dense preference over short answers (chat/style/refusals); PPO/RLHF for sparse shaped long-horizon rewards (reasoning, tools).
  6. 06“Measured results?” → InstructGPT 1.3B > 175B; 85±3% preferred over GPT-3; DPO ~61% vs PPO ~57% win rate on TL;DR.
  7. 07“How does alignment regress capability?” → the alignment tax: RLHF narrows diversity (QDC), preference tuning can drop MMLU/SQuAD; gate on a general eval.
  8. 08“DPO vs PPO failures?” → reward hacking (PPO) vs preference warping + OOD brittleness against a frozen reference (DPO).
Going deeper: “what is iterative/online DPO?” (sample best-of-N from the current policy, label with a reward model or self-judge, retrain — online RLHF without PPO); “name DPO’s cousins” (KTO, IPO, ORPO, SimPO — same loss family, different data formats / no-reference variants); “your preference data has low annotator agreement” (kappa, escalate, drop — the same data discipline as L2); and “why do reasoning models use RL not DPO?” (sparse, verifiable, long-horizon reward — RL propagates credit; pure DPO’s dense pairwise signal doesn’t). Tie each answer to the reward-signal structure.

Could you derive DPO from the RLHF objective, choose DPO vs PPO by the reward signal, and explain the alignment tax + how the two methods fail differently?

New to itGetting thereConfident

Takeaways

  • RLHF = SFT → reward model (Bradley-Terry) → PPO with a KL penalty; failures are reward hacking, instability, and four-model cost.
  • DPO collapses that into one supervised loss: the reward is implicit in the policy/reference log-ratio — no RM, no PPO sampling.
  • Choose by signal structure: DPO for dense short-answer preference (chat/style); PPO/RLHF for sparse shaped long-horizon rewards (reasoning/tools).
  • Numbers: InstructGPT 1.3B preferred over 175B (85±3% over GPT-3); DPO ~61% vs PPO ~57% on TL;DR.
  • Alignment tax: RLHF narrows diversity (QDC); preference tuning can regress MMLU/SQuAD — gate on a general eval.
  • Failures differ in kind: PPO reward-hacks an explicit RM; DPO preference-warps and is brittle OOD against a frozen reference.

Next: quantization & efficient serving — GPTQ/AWQ/fp8/GGUF tradeoffs, KV-cache reuse, continuous batching, speculative decoding, and the serving stacks.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.