Distillation transfers a larger "teacher" model's behavior into a smaller "student" model. Two main modes: (1) data distillation: sample completions from the teacher on a large prompt set, then SFT the student on (prompt, teacher_response) pairs; (2) logits distillation: match the teacher's full softmax distribution (which encodes "dark knowledge" about relative confidences), typically via KL-divergence between teacher and student logits at training time. Distillation can match 80-95% of teacher quality at 5-10x smaller model size, but real-world gains depend heavily on the data distribution match to deployment. Trade-offs: synthetic-data distillation generalizes better than just SFT-imitating human answers because teacher confusion on hard examples is implicitly conveyed. Senior nuance: classic KD vs sequence-level KD vs multi-teacher ensembling; also "self-distillation" can sometimes improve a model without a smaller student. Distillation is also used in RLHF to bootstrap reward models.