Present the three-stage pipeline, then go deep where the interviewer steers.
Stage 1 — SFT: curate high-quality demonstrations (human-written + filtered synthetic); train the base model to follow instructions. Quality >> quantity; the data mixture is a first-class design decision.
Stage 2 — Reward model: collect pairwise preferences (annotators pick the better of two responses — far more reliable than absolute scores); train an RM (usually the SFT model + scalar head) on Bradley-Terry loss. The data operation is the hard part: rater guidelines, inter-annotator agreement tracking, disagreement adjudication, rater-pool diversity — say this; labs grade for it.
Stage 3 — RL: PPO against the RM with a KL penalty to the SFT policy — the KL term is what stops the policy from wandering into RM-exploiting gibberish. Infra reality: four models in memory (policy, reference, RM, value) — this is why RLHF infra is expensive and why the alternative exists.
DPO vs PPO — the expected discussion: DPO optimizes preferences directly with a classification-style loss (no rollouts, dramatically simpler infra, more stable); PPO enables online exploration and iterated data collection and still edges it for frontier quality. Reasonable answer: DPO first for iteration speed, PPO/online methods when you have the infra and need the last few points.
The central failure mode — reward hacking: the policy exploits RM blind spots (length bias → verbosity, sycophancy, confident hedging). Mitigations: KL constraint, RM ensembles, iterated RM retraining on fresh policy samples, and held-out human eval as the final arbiter — never trust RM score alone; it's Goodhart's law in production.
Evaluation: win-rate vs the SFT baseline (human + LLM judge), capability regression suites (alignment-tax check), safety red-teaming.
Follow-up probes: RM score goes up, human eval goes down — diagnose. (Reward hacking; inspect high-RM samples, retrain the RM on adversarial pairs.) Where do RLAIF / constitutional approaches slot in? Online (PPO/GRPO) vs offline (DPO) preference optimization — sample efficiency and distribution-shift tradeoffs. How does this differ for reasoning models? (Verifiable rewards — unit tests, math checkers — replace learned RMs where outputs are checkable; RLVR.)