Greedy decoding (always pick the argmax) is deterministic but loops and boring. Sampling injects diversity by drawing from a probability distribution. Temperature scales the logits before softmax: low T sharpens the distribution toward greedy, high T flattens it. Top-k truncates to only the k highest-probability tokens before sampling, eliminating long-tail noise but capping diversity rigidly. Top-p (nucleus) samples from the smallest set of tokens whose cumulative probability mass exceeds p, which adapts to the entropy of the distribution: tight distributions yield tiny pools, flat ones larger. For code/data extraction: low temperature (0-0.2), top-p=1, often top-k off. For open-ended creative writing: temperature ~0.7-0.9, top-p ~0.9, top-k 40-100. For chat assistants: temperature 0.6-0.8, top-p ~0.9. Senior nuance: temperature and top-p are NOT redundant; they operate in different spaces (logit scale vs probability mass), and combining them lets you tune shape-of-distribution AND cut-off separately. Beam search is sometimes better for narrow NLG tasks (summarization) where argmax-quality matters.