After SFT teaches format, preference tuning teaches taste — the RLHF three-stage pipeline (SFT → reward model → PPO) and its failure modes, DPO’s closed-form collapse of that pipeline into one loss, when each wins (dense preferences vs sparse shaped rewards), and how alignment regresses capabilities and diversity.
SFT (L2) makes a base model follow instructions, but it can’t easily express “this answer is better than that one” — it only sees positive examples. Preference optimization is the stage that aligns the model to human preference over pairs: helpfulness, tone, refusal behaviour, the difference between an acceptable answer and a great one. The headline result — InstructGPT’s 1.3B aligned model preferred over 175B GPT-3 — comes from this stage. The modern debate is RLHF (the original three-stage pipeline) vs DPO (one supervised loss), and getting the tradeoff right is a senior alignment interview’s core.
The RLHF pipeline (Ouyang et al., 2022) is three stages, and a strong answer walks all three. (1) SFT: fine-tune the base model on curated demonstrations to get a reference policy. (2) Reward model (RM): train a scalar model r(x,y) on pairwise human preference data {y_w, y_l} (winner/loser) with a Bradley-Terry loss -log sigmoid(r(x,y_w) - r(x,y_l)) — the RM learns to score any (prompt, response). (3) PPO: optimise the policy with reinforcement learning against the RM, plus a KL penalty toward the SFT reference so the policy doesn’t drift into gibberish that games the RM. InstructGPT used ~33k prompts for the RM stage and ~31k for PPO, with a pretraining-mix term (γ ≈ 27.8) to fight regression on public NLP benchmarks.
code
1RLHF (InstructGPT) -- three stages, four models live in memory23 1. SFT base -> reference policy (demonstrations)4 2. REWARD MODEL r(x,y) scalar (pairwise prefs, Bradley-Terry)5 loss = -log sigmoid( r(x,y_w) - r(x,y_l) )6 3. PPO maximise r(x,y) - beta·KL(policy || SFT_ref)7 (policy + reference + reward + value models all resident)89 Failure modes: reward hacking (game the RM), PPO instability / mode collapse,10 4-model infra cost. KL penalty + pretraining-mix (gamma~27.8) fight drift.
The PPO failure modes are the part interviewers probe hardest. Reward hacking: the policy finds outputs that score high on the RM but are degenerate to humans (the RM is a proxy, and Goodhart’s law applies — “reward over-optimisation”). Instability: PPO can collapse to a narrow, high-reward, low-diversity mode. Infrastructure cost: you keep four models in memory — policy, reference, reward, and value — which is why RLHF is heavy to run. Interview angle. A weak answer stops at “RLHF aligns the model.” A strong one names reward hacking vs PPO instability vs the 4-model cost, and notes the KL penalty is the lever that keeps the policy near the SFT reference.
DPO — your language model is secretly a reward model
DPO (Rafailov et al., 2023) starts from a clever observation: under the Bradley-Terry model and a KL-bounded objective, the optimal RLHF policy can be written in closed form as a function of the reference policy and the reward — which means the reward model and the policy aren’t separate objects. The reward is implicit in the log-ratio of the policy to the reference. So DPO eliminates the reward model and the PPO loop entirely, collapsing all three stages after SFT into a single supervised classification loss over (chosen, rejected) pairs. The paper’s subtitle says it: “your language model is secretly a reward model.”
code
1DPO -- one supervised loss, no reward model, no PPO sampling23 L_DPO = -log sigmoid( beta · [ log( pi(y_w|x) / pi_ref(y_w|x) )4 - log( pi(y_l|x) / pi_ref(y_l|x) ) ] )56 pi = the policy being trained7 pi_ref = the FROZEN SFT model (reference)8 beta = KL strength (default 0.1; 0.5 for TL;DR summarization)9 y_w/y_l = chosen / rejected response1011 No reward model. No sampling from the policy during training. No PPO.12 It's a Bradley-Terry classification loss on preference pairs.
The practical consequences are why DPO took over open-weight chat alignment. There is no reward model to train, no sampling from the policy during training (PPO’s expensive, unstable rollouts are gone), and no PPO infrastructure. The paper reports DPO is “stable, performant, and computationally lightweight” and removes “the need for … significant hyperparameter tuning.” The one finicky knob is beta (the KL strength, default 0.1, 0.5 for TL;DR summarisation): lower beta lets the policy diverge more from the reference. The canonical recipe — load a 4-bit base, SFT on instructions, then DPO with beta=0.1 and a small learning rate (~1e-6) on preference pairs — is the de-facto standard for small-team and open-weight alignment.
When each wins — dense preferences vs sparse shaped rewards
The operational division mlabonne’s course states cleanly: DPO is “great for creating chat models”; PPO/RL is “ideal for creating reasoning models.” The mechanism explains why. DPO excels when the target is a dense preference over short responses — style, tone, refusal patterns — because the loss directly prefers one answer over another and the chosen/rejected signal is rich. PPO/RLHF excels when the optimisation signal is sparse and reward-shaped — math correctness, multi-step reasoning, tool-call success — because a reward model can propagate credit across a long horizon in ways a binary preference loss cannot. This is also why modern reasoning models lean on RL (and variants like GRPO) rather than pure DPO.
code
1WHICH PREFERENCE METHOD? (match the reward signal's structure)23 Signal shape -> Method4 ---------------------------------------- ----------------------------5 dense preference over short answers -> DPO (chat, tone, refusals)6 (style / helpfulness / format taste) stable, 1 loss, no RM78 sparse, shaped, long-horizon reward -> PPO / RLHF (reasoning, tools,9 (math correctness, multi-step, tools) math) -- RM propagates credit1011 Calibrated reward signal needed for -> RLHF (you get a reusable RM12 online iteration / continuous eval for eval + online loops)1314 Most open chat models: SFT -> DPO. Reasoning models: RL (PPO/GRPO).
The measured results bound the debate. InstructGPT: 175B InstructGPT outputs preferred over 175B GPT-3 85±3% of the time, and over few-shot 175B GPT-3 71±4%. DPO: it “exceeds PPO-based RLHF in ability to control sentiment” on IMDb, and on TL;DR summarisation hits a win rate of ~61% at temperature 0 versus PPO’s ~57% at its best sampling temperature. So DPO isn’t just cheaper — on these preference-shaped tasks it’s competitive or better. Interview angle. “When would you still reach for PPO/RLHF in 2026?” → when the signal is sparse/shaped (reasoning, tool success), when you need a calibrated reward model for online iteration and continuous eval, or when you’re running RLHF steps during deployment — not “PPO is the standard.”
How preference tuning regresses capabilities
Alignment is not free, and naming its costs is the senior move. The QDC study (L2) found that RLHF models generalise better than SFT models but exhibit worse sample diversity — alignment improves in-distribution accuracy while narrowing the output distribution. This is the “alignment tax”: a heavily preference-tuned model can become repetitive, over-hedge, or lose creative range. InstructGPT explicitly fought regression on public NLP benchmarks (SQuADv2, DROP) with a pretraining-mix term and tested KL coefficients up to 100× the default — and still reported residual regressions. The lesson: preference tuning trades expressivity for adherence, and you must measure the trade, not assume it’s positive.
DPO has its own characteristic failure, distinct in kind from PPO’s reward hacking: preference warping (and reference brittleness). Because DPO optimises against a frozen reference and a fixed set of preference pairs, it can overfit to the shape of those pairs rather than true quality — producing outputs that are stylistically aligned but factually shallow. And if the SFT reference drifts (you change it) or the pairs go stale, the DPO signal degrades; the paper itself flags out-of-distribution generalisation and reward over-optimisation in the DPO setting as open limitations. Interview angle. “How does DPO fail, and how is that different from PPO?” → PPO fails by reward hacking (gaming an explicit RM); DPO fails by preference warping and OOD brittleness against a frozen reference. Knowing the failures differ in kind is the reading-depth signal.
code
1THE ALIGNMENT TAX (preference tuning costs something)23 Symptom Cause Mitigation4 ----------------------------- ------------------------- ----------------5 repetitive / low-diversity RLHF narrows output dist. sampling, rejection-6 output (QDC: better gen, worse D) sampling, seed diversity7 general benchmark regression over-optimising the reward pretraining-mix term,8 (MMLU/SQuAD/DROP drop) / capability drift KL penalty, gate on eval9 stylistically aligned but DPO preference warping diverse, fresh prefs;10 factually shallow (overfits pair shape) watch OOD; don't over-train11 signal goes stale frozen reference drifts re-pair if you move the SFT ref1213 Always run a GENERAL eval alongside the preference eval to catch the tax.
The cross-cutting insight that ties the alignment lesson back to the data lesson: both pipelines degrade if the SFT data/reference is stale, so the lifecycle of the SFT corpus — not the choice of DPO vs PPO — is often the highest-leverage operational decision. A drifting reference quietly poisons DPO; a weak SFT stage caps everything downstream in RLHF. This is the same theme as the rest of the track: the data and the reference are the load-bearing parts, and the algorithm is the smaller decision on top.
You’re aligning a customer-support chat model to a preferred tone and refusal style, with a small team and limited GPU budget. Which approach fits best and why?
AFull RLHF with PPO and a trained reward modelBSFT then DPO — a dense style/refusal preference is exactly DPO’s sweet spot: one stable loss, no reward model, no PPO infraCPrompt engineering only — never tune preferences
In an interview you explain DPO. What’s the one-sentence core that signals real understanding?
ADPO doesn’t need a reward modelBUnder Bradley-Terry + a KL-bounded objective, the optimal policy is a closed-form function of the reference and reward, so the reward is implicit in the policy/reference log-ratio — letting DPO replace the RM and PPO with one classification lossCDPO uses reinforcement learning more efficiently than PPO
You’re building a model that must solve multi-step math and verify tool-call success. Why might PPO/RLHF beat DPO here?
APPO trains faster than DPOBThe reward is sparse and shaped (correctness over long horizons); a reward model can propagate credit across steps in ways a binary preference loss can’tCDPO can’t be used on math problems at all
After heavy DPO the model is more on-tone but its answers feel repetitive and slightly shallower factually. What’s happening and what do you do?
ANothing — preference tuning only improves the modelBIt’s the alignment tax (reduced diversity + DPO preference warping); diversify/refresh preference pairs, don’t over-train, adjust sampling, and gate on a general eval to bound the regressionCRaise beta to 1.0 to make the model more confident
Which statement about RLHF failure modes vs DPO failure modes is correct?
ABoth fail identically — there’s no meaningful differenceBPPO/RLHF fails by reward hacking (gaming an explicit reward model); DPO fails by preference warping and OOD brittleness against a frozen referenceCDPO is immune to over-optimisation because it has no reward model
Alignment rounds test whether you can derive DPO (not just recite “no reward model”), walk the RLHF three stages with their failure modes, choose between them by the reward signal’s structure, and name the alignment tax. DPO is the quiet consensus default for chat-style alignment; the senior signal is knowing precisely when PPO/RLHF still earns its cost — and that the two methods fail in different ways.
01“Explain RLHF end-to-end.” → SFT → reward model (Bradley-Terry pairwise) → PPO with a KL penalty to the SFT reference; four models resident.
02“PPO failure modes?” → reward hacking (game the RM, Goodhart), instability/mode collapse, and the four-model infra cost.
03“How does DPO change it?” → closed-form optimal policy under BT+KL ⇒ reward implicit in the policy/reference log-ratio ⇒ one classification loss, no RM, no PPO.
04“DPO’s finicky knob?” → beta (KL strength), default 0.1 (0.5 for TL;DR); lower beta = more divergence from the reference.
05“When DPO vs PPO?” → DPO for dense preference over short answers (chat/style/refusals); PPO/RLHF for sparse shaped long-horizon rewards (reasoning, tools).
06“Measured results?” → InstructGPT 1.3B > 175B; 85±3% preferred over GPT-3; DPO ~61% vs PPO ~57% win rate on TL;DR.
07“How does alignment regress capability?” → the alignment tax: RLHF narrows diversity (QDC), preference tuning can drop MMLU/SQuAD; gate on a general eval.
08“DPO vs PPO failures?” → reward hacking (PPO) vs preference warping + OOD brittleness against a frozen reference (DPO).
Going deeper: “what is iterative/online DPO?” (sample best-of-N from the current policy, label with a reward model or self-judge, retrain — online RLHF without PPO); “name DPO’s cousins” (KTO, IPO, ORPO, SimPO — same loss family, different data formats / no-reference variants); “your preference data has low annotator agreement” (kappa, escalate, drop — the same data discipline as L2); and “why do reasoning models use RL not DPO?” (sparse, verifiable, long-horizon reward — RL propagates credit; pure DPO’s dense pairwise signal doesn’t). Tie each answer to the reward-signal structure.
Could you derive DPO from the RLHF objective, choose DPO vs PPO by the reward signal, and explain the alignment tax + how the two methods fail differently?
New to itGetting thereConfident
Takeaways
RLHF = SFT → reward model (Bradley-Terry) → PPO with a KL penalty; failures are reward hacking, instability, and four-model cost.
DPO collapses that into one supervised loss: the reward is implicit in the policy/reference log-ratio — no RM, no PPO sampling.
Choose by signal structure: DPO for dense short-answer preference (chat/style); PPO/RLHF for sparse shaped long-horizon rewards (reasoning/tools).
Numbers: InstructGPT 1.3B preferred over 175B (85±3% over GPT-3); DPO ~61% vs PPO ~57% on TL;DR.
Alignment tax: RLHF narrows diversity (QDC); preference tuning can regress MMLU/SQuAD — gate on a general eval.
Failures differ in kind: PPO reward-hacks an explicit RM; DPO preference-warps and is brittle OOD against a frozen reference.
Next: quantization & efficient serving — GPTQ/AWQ/fp8/GGUF tradeoffs, KV-cache reuse, continuous batching, speculative decoding, and the serving stacks.