Overfitting = model captures noise/idiosyncrasies in training data, fails to generalize. Detected by: gap between training and validation/test error that grows over training; poor performance on holdout set. Prevention layers: (1) data: more data, better data quality, data augmentation; (2) model: simpler architecture, fewer features; (3) regularization: L1/L2, dropout, early stopping, weight decay; (4) ensemble: bagging (random forests) reduces variance; (5) cross-validation: k-fold for robust estimate of generalization error. Senior nuance: in deep learning the "double descent" phenomenon complicates the overfitting story: some overparameterized models generalize better than smaller ones (interpolation regime). Production-specific: monitoring the train-val gap over time catches overfitting retraining pipelines. Don't conflate train-test gap with real-world shift.