SFT (supervised fine-tuning) teaches a base model to imitate demonstrations, but it can only learn what's already in the data and cannot directly optimize for "what users actually prefer" when there are many equally valid completions. RLHF adds a preference-modeling step: humans rank outputs (often pairwise), a reward model r(x, y) is trained on those preferences, and the LM is then optimized against that reward signal. Two flavors: (1) classic RLHF uses PPO with a KL penalty to the reference model; (2) Direct Preference Optimization (DPO) re-derives the same objective as a supervised loss on preference pairs: simpler, more stable, no reward model needed, dominates practical pipelines by 2025. Conclusion: RLHF/DPO is what makes ChatGPT-style assistants possible; it shapes style, refusal behaviors, helpfulness, harmlessness. Senior nuances: reward hacking, distribution shift between RM and policy, "alignment tax", and the emergence of RLHF-V, RLHF-Diffusion, online DPO, GRPO for reasoning models in 2025-2026.