Lesson 3 of 6 · 49 min
Regularization & calibration
The variance dial: L1/L2/elastic-net geometry and when each wins, dropout’s implicit ensemble, early stopping as spectral regularization. Then the orthogonal problem — why probabilities miscalibrate, Platt vs isotonic (the 1000-sample rule), and ECE/reliability diagrams.
Two different fixes people constantly conflate
J_λ(θ) = J(θ) + λ·R(θ), where “λ ≥ 0 balances fitting the data (small J) against small model complexity (small R)” (CS229). Two endpoints anchor intuition: λ=0 recovers the unregularized loss; “an extremely large λ can result in a model with large bias.” So regularization is literally a knob on the Lesson-1 tradeoff — you are buying variance reduction with bias, and the art is the smallest λ that closes the generalization gap. Google’s MLCC names λ the “regularization rate” and notes L2 “encourages weights toward 0 but never all the way to zero” — the precise contrast with L1 that interviewers fish for.L1 vs L2 vs elastic-net: the geometry is the answer
Σwⱼ² — a spherical constraint region; the loss contour touches it in the interior, so weights shrink smoothly but stay non-zero. L1 (lasso) penalizes Σ|wⱼ| — a diamond whose corners sit on the axes; the optimum tends to land on a corner where many weights are exactly zero, giving sparsity / feature selection. ESL’s precise phrasing: lasso is “the smallest q such that the constraint region is convex,” and its non-differentiable corners are exactly why solutions are sparse.1L1 vs L2 vs ELASTIC-NET23 penalty region effect on weights use when4 ------------ ------------ -------------------- ----------------------------5 L2 (ridge) sphere shrink all, none = 0 correlated/ill-conditioned X6 L1 (lasso) diamond many exactly 0 want sparsity / selection7 elastic-net rounded diamond mix of both sparse AND correlated (genomics)89 Correlated features: lasso picks ONE arbitrarily and zeros the rest;10 ridge shrinks the whole group TOGETHER;11 elastic-net selects a group and shrinks within it.λ·Σ(α·wⱼ² + (1−α)|wⱼ|) — “selects variables like the lasso and shrinks together the coefficients of correlated predictors like ridge.” Interview angle. “Explain ridge vs lasso to a non-technical stakeholder” has a memorized one-liner: ridge = “keep all features but reduce their impact”; lasso = “drop the features that contribute little.” And “why is L2 the default in neural nets, L1 rare?” → L1’s non-smoothness doesn’t pair well with SGD’s gradient flow, whereas L2 is weight decay and integrates cleanly.
Ridge vs Lasso Regression, Visualized!!!StatQuest with Josh StarmerL2 shrinks; L1 selects. Everything else about ridge versus lasso is downstream of that one sentence — so answer the geometry first, then let the feature statistics pick the side.
l1_ratio α) is itself a hyperparameter: α=1 is pure lasso, α=0 pure ridge, and the genomics-style use case — thousands of correlated features where you want both selection and grouped shrinkage — typically lands around α≈0.5, tuned by CV alongside λ. Second, always standardize features before any penalty: L1/L2 penalize coefficient magnitude, so an unscaled feature in different units gets penalized arbitrarily. Forgetting to scale before regularizing is a classic silent bug — and a question interviewers slip in to see if you have actually run this.Dropout & early stopping: regularizers that aren’t penalties
p; at inference you keep all units but scale activations by (1−p) (or, equivalently, “inverted dropout” scales up by 1/(1−p) during training and does nothing at test) so the expected input to each unit matches. Forgetting to switch off dropout at inference — e.g. leaving the model in train() mode in PyTorch — is a real, common bug that silently degrades predictions. The expectation-matching is exactly what makes the test-time network approximate the trained ensemble’s average.Key idea
patience on a validation metric you are already computing. On any iterative learner (GBT, neural net) it is the first thing to turn on and the last thing to turn off.Key idea
Calibration: a separate step, not a property of the loss
p, the empirical positive rate is p. This is a property of the probabilities, not of the argmax decision — and it is orthogonal to discrimination (AUC): you can have a high-AUC model that is badly calibrated, and calibration leaves AUC unchanged. The directions of miscalibration are predictable (Niculescu-Mizil & Caruana): boosted trees and SVMs “push probability mass away from 0 and 1,” producing a characteristic sigmoid distortion; naive Bayes pushes toward 0/1; and modern deep nets are typically over-confident (Guo et al., 2017) — the reliability-diagram bars sit below the diagonal.P(y=1|f) = σ(A·f + B) — best when the distortion is sigmoid-shaped and the calibration set is modest. Isotonic regression fits a non-parametric monotone step function — “can correct any monotonic distortion” but is “more prone to overfitting.” The memorized rule (Niculescu-Mizil & Caruana): Platt “performs better than isotonic for small-to-medium calibration sets (< 1000 cases)”; with “1000 or more points, isotonic always yields performance as good as or better than Platt.” For deep-net logits, temperature scaling (a single T dividing the logits) is the standard one-knob fix.1CALIBRATION CHEAT-SHEET23 method fits best when4 ------------- ---------------------- ------------------------------------5 Platt scaling sigmoid(A*f + B) sigmoid distortion, < 1000 cal pts6 Isotonic monotone step function any monotone distortion, >= 1000 pts7 Temperature logits / T (one scalar) deep-net logits (keeps argmax fixed)89 ECE = sum_m (|B_m|/n) * | acc(B_m) - conf(B_m) | (M bins, usually 10-15)1011 Reliability diagram: bar BELOW diagonal = over-confident (conf > acc);12 bar ABOVE diagonal = under-confident.13 ALWAYS measure ECE on a held-out calibration fold BEFORE you calibrate.ECE = Σ_m (|B_m|/n)·|acc(B_m) − conf(B_m)|, a bin-weighted average gap between accuracy and confidence. The senior caveat: ECE is binning-dependent (10–15 bins is standard) — quoting an absolute ECE without the binning scheme is meaningless, and variants like ACE/TACE use adaptive bins. The publishable bar for top-confidence buckets is roughly ECE < 0.01.
KDD 2020 Tutorial: How to Calibrate Your Neural Network ClassifierACMKey idea
When calibration matters — and when it is wasted effort
CalibratedClassifierCV does this with an internal CV split), because calibrating on the training data just re-learns the same over-confidence. And calibration is monotone — it re-maps scores without reordering them — so it changes Brier/ECE but leaves AUC and the ranking essentially fixed. That is the whole reason it is a safe, cheap post-processing step: you can always add it to make probabilities honest without risking your discrimination metric.Interview prep
- 01“L1 vs L2?” → L1 (diamond) zeros weights → sparsity/selection; L2 (sphere) shrinks all → handles correlated/ill-conditioned features.
- 02“What if features are correlated?” → lasso keeps one arbitrarily; ridge shrinks the group together; elastic-net selects-and-shrinks.
- 03“Why L2 (weight decay) in nets, not L1?” → L1’s non-smoothness fights SGD; L2 integrates cleanly with gradient flow.
- 04“Early stopping — is it real regularization?” → yes, spectral: halting iteration constrains complexity; it’s the free regularizer on any val-fold run.
- 05“Is calibration the same as accuracy/AUC?” → no — it’s about probability quality and is orthogonal to ranking; calibrating leaves AUC unchanged.
- 06“Platt vs isotonic?” → Platt for sigmoid distortion and < 1000 cal points; isotonic for ≥ 1000 (any monotone distortion, but overfits when small).
- 07“How do you measure miscalibration?” → reliability diagram + ECE = bin-weighted |accuracy − confidence|; report the binning scheme.
- 08“When does calibration actually matter?” → moving thresholds, cost-sensitive decisions, model blending, risk communication — not fixed-threshold ranking.
Common mistake
The #1 red-flag answer: “Calibrate the model to improve its accuracy / AUC.”
Checkpoint
You have ~2,000 features, strongly suspect only a few dozen carry signal, and the business wants a short, explainable feature list. Which regularizer and why?
Checkpoint
A gradient-boosted fraud model has 0.94 ROC-AUC, but its predicted probabilities are needed to set a cost-based threshold that will shift weekly. A reliability diagram shows bars well below the diagonal. Best step?
Checkpoint
A model produces a group of three highly correlated features. Lasso keeps one and zeros the other two, and the chosen one flips between resamples. What is the most stable fix that still controls variance?
Checkpoint
You have only ~400 examples in your held-out calibration set and need to calibrate a classifier whose distortion looks sigmoidal. Which method, and what’s the risk of the alternative?
Checkpoint
A teammate adds L2, dropout, AND aggressive early stopping to a model that was only mildly overfitting; now it underfits. What’s the principled correction?
Could you pick L1/L2/elastic-net from feature statistics, explain why dropout/early-stopping regularize, and decide Platt-vs-isotonic and whether calibration even matters?
Takeaways
- Regularization is the variance dial: L1 (diamond) → sparsity/selection; L2 (sphere) → smooth shrinkage for correlated features; elastic-net mixes both.
- Dropout ≈ averaging an exponential ensemble (MNIST 1.6→1.35%); early stopping is spectral regularization and free on any val fold.
- Match one regularizer to the prior; stacking them blindly just adds bias.
- Calibration is orthogonal to AUC: it fixes probability quality (ECE/Brier), not ranking.
- Platt for < 1000 calibration points / sigmoid distortion; isotonic for ≥ 1000; temperature for deep-net logits — and only calibrate when a decision consumes the probability.
Next: model selection in practice — logistic vs gradient-boosted trees vs neural nets vs LLMs, and when to use which.
Sources
Free to read · better with Enzo
Learn it with Enzo
Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.