Classic RLHF (Ouyang et al. 2022) uses PPO with reward model + KL constraint. Problems: reward hacking, instability, on-policy rollout cost. DPO (2023) closed-formed the same objective as supervised loss, eliminating the reward model and RL loop. IPO (2024) softened DPO's Bradley-Terry assumption for noisier preferences. KTO (2024) used single-rating (good/bad) instead of pairs: cheaper labeling. Online DPO (2024) iterates between sampling from current policy and updating on freshly ranked pairs. GRPO (DeepSeek-R1, early 2025) replaced the value head with group-relative advantages: sample N completions per prompt, normalize their rewards, use normalized advantages as update signal, dramatically cheaper than PPO for reasoning models. Process-reward models (PRMs) reward intermediate reasoning steps and emerged alongside GRPO. Senior view: GRPO is the dominant 2025-2026 recipe for reasoning fine-tunes; classic PPO is mostly used in safety/RLHF contexts; DPO remains default for chat-style alignment.