Lesson 3 of 6 · 49 min

Regularization & calibration

The variance dial: L1/L2/elastic-net geometry and when each wins, dropout’s implicit ensemble, early stopping as spectral regularization. Then the orthogonal problem — why probabilities miscalibrate, Platt vs isotonic (the 1000-sample rule), and ECE/reliability diagrams.

Two different fixes people constantly conflate

Regularization and calibration both make models “better,” but they fix different things. Regularization turns the variance dial down at a bias cost — it changes what the model is. Calibration is a post-hoc audit of whether the model’s probabilities mean what they say — it leaves ranking (AUC) untouched and only fixes the numbers. The interview tests both, and tests whether you know they are orthogonal: “L1 vs L2 — and what if features are correlated?”, “is your model calibrated, and does it matter here?”, “Platt or isotonic?” Conflating them — or calibrating by default — is the red flag.
The penalized objective is J_λ(θ) = J(θ) + λ·R(θ), where “λ ≥ 0 balances fitting the data (small J) against small model complexity (small R)” (CS229). Two endpoints anchor intuition: λ=0 recovers the unregularized loss; “an extremely large λ can result in a model with large bias.” So regularization is literally a knob on the Lesson-1 tradeoff — you are buying variance reduction with bias, and the art is the smallest λ that closes the generalization gap. Google’s MLCC names λ the “regularization rate” and notes L2 “encourages weights toward 0 but never all the way to zero” — the precise contrast with L1 that interviewers fish for.

L1 vs L2 vs elastic-net: the geometry is the answer

Answer L1-vs-L2 with the constraint-region picture before any equation (Prepaired, Amit Maurya: “strong answers are scenario-driven; weak answers are equation-driven”). L2 (ridge) penalizes Σwⱼ² — a spherical constraint region; the loss contour touches it in the interior, so weights shrink smoothly but stay non-zero. L1 (lasso) penalizes Σ|wⱼ| — a diamond whose corners sit on the axes; the optimum tends to land on a corner where many weights are exactly zero, giving sparsity / feature selection. ESL’s precise phrasing: lasso is “the smallest q such that the constraint region is convex,” and its non-differentiable corners are exactly why solutions are sparse.
code
1L1 vs L2 vs ELASTIC-NET23  penalty        region         effect on weights      use when4  ------------   ------------   --------------------   ----------------------------5  L2 (ridge)     sphere         shrink all, none = 0   correlated/ill-conditioned X6  L1 (lasso)     diamond        many exactly 0         want sparsity / selection7  elastic-net    rounded diamond mix of both           sparse AND correlated (genomics)89  Correlated features:  lasso picks ONE arbitrarily and zeros the rest;10                        ridge shrinks the whole group TOGETHER;11                        elastic-net selects a group and shrinks within it.
The correlated-features follow-up is “almost always asked” (Priyaa). The crisp answer: with a cluster of correlated predictors, lasso arbitrarily keeps one and zeros the rest (unstable across resamples), ridge shrinks them together (stable but dense), and elastic-net — penalty λ·Σ(α·wⱼ² + (1−α)|wⱼ|) — “selects variables like the lasso and shrinks together the coefficients of correlated predictors like ridge.” Interview angle. “Explain ridge vs lasso to a non-technical stakeholder” has a memorized one-liner: ridge = “keep all features but reduce their impact”; lasso = “drop the features that contribute little.” And “why is L2 the default in neural nets, L1 rare?” → L1’s non-smoothness doesn’t pair well with SGD’s gradient flow, whereas L2 is weight decay and integrates cleanly.
Ridge vs Lasso Regression, Visualized!!!StatQuest with Josh Starmer
L2 shrinks; L1 selects. Everything else about ridge versus lasso is downstream of that one sentence — so answer the geometry first, then let the feature statistics pick the side.
Two practitioner notes that turn the theory into a recipe. First, elastic-net’s mixing parameter (scikit-learn’s l1_ratio α) is itself a hyperparameter: α=1 is pure lasso, α=0 pure ridge, and the genomics-style use case — thousands of correlated features where you want both selection and grouped shrinkage — typically lands around α≈0.5, tuned by CV alongside λ. Second, always standardize features before any penalty: L1/L2 penalize coefficient magnitude, so an unscaled feature in different units gets penalized arbitrarily. Forgetting to scale before regularizing is a classic silent bug — and a question interviewers slip in to see if you have actually run this.

Dropout & early stopping: regularizers that aren’t penalties

Dropout (Srivastava et al., 2014) randomly zeros units during training and, at test time, uses all units with scaled weights — a Monte-Carlo approximation to averaging an exponential ensemble of thinned sub-networks. The measured impact is concrete: MNIST MLP error 1.6% → 1.35%; CIFAR-10 14.98% → 12.61% (dropout on every layer); TIMIT phone-error 23.4% → 21.8%. The authors note dropout “prevents overfitting” and “does not even need early stopping” — a deliberate antagonist point: do not stack redundant regularizers without a reason.
One dropout subtlety interviewers probe: train vs test behavior differs. During training you randomly drop units with probability p; at inference you keep all units but scale activations by (1−p) (or, equivalently, “inverted dropout” scales up by 1/(1−p) during training and does nothing at test) so the expected input to each unit matches. Forgetting to switch off dropout at inference — e.g. leaving the model in train() mode in PyTorch — is a real, common bug that silently degrades predictions. The expectation-matching is exactly what makes the test-time network approximate the trained ensemble’s average.
Early stopping is the cheapest regularizer you will ever use: split off a validation set and “stop training as soon as the error on the validation set is higher than it was the last time it was checked” (Prechelt). It is not a hack — it is spectral regularization: halting the iteration imposes a smoothness/low-complexity constraint, yielding a bias-variance decomposition in the iteration index itself. Prechelt frames the choice of stopping criterion as an explicit time-vs-generalization tradeoff (~1% test-error improvement can cost ~4× training time). Interview angle. “Name four regularizers and when each.” → L2 for dense/correlated features; L1 when you want sparse explanations; dropout for dense deep nets; early stopping on every iterative run that has a validation fold (free).

Calibration: a separate step, not a property of the loss

A classifier is calibrated if, among all instances it scores p, the empirical positive rate is p. This is a property of the probabilities, not of the argmax decision — and it is orthogonal to discrimination (AUC): you can have a high-AUC model that is badly calibrated, and calibration leaves AUC unchanged. The directions of miscalibration are predictable (Niculescu-Mizil & Caruana): boosted trees and SVMs “push probability mass away from 0 and 1,” producing a characteristic sigmoid distortion; naive Bayes pushes toward 0/1; and modern deep nets are typically over-confident (Guo et al., 2017) — the reliability-diagram bars sit below the diagonal.
Two fixes, and the rule that decides between them is the single most-asked calibration question. Platt scaling fits a one-parameter logistic on the scores, P(y=1|f) = σ(A·f + B) — best when the distortion is sigmoid-shaped and the calibration set is modest. Isotonic regression fits a non-parametric monotone step function — “can correct any monotonic distortion” but is “more prone to overfitting.” The memorized rule (Niculescu-Mizil & Caruana): Platt “performs better than isotonic for small-to-medium calibration sets (< 1000 cases)”; with “1000 or more points, isotonic always yields performance as good as or better than Platt.” For deep-net logits, temperature scaling (a single T dividing the logits) is the standard one-knob fix.
code
1CALIBRATION CHEAT-SHEET23  method          fits                     best when4  -------------   ----------------------   ------------------------------------5  Platt scaling   sigmoid(A*f + B)         sigmoid distortion, < 1000 cal pts6  Isotonic        monotone step function   any monotone distortion, >= 1000 pts7  Temperature     logits / T (one scalar)  deep-net logits (keeps argmax fixed)89  ECE = sum_m (|B_m|/n) * | acc(B_m) - conf(B_m) |     (M bins, usually 10-15)1011  Reliability diagram: bar BELOW diagonal = over-confident (conf > acc);12                       bar ABOVE diagonal = under-confident.13  ALWAYS measure ECE on a held-out calibration fold BEFORE you calibrate.
Expected Calibration Error is the scalar summary of a reliability diagram: ECE = Σ_m (|B_m|/n)·|acc(B_m) − conf(B_m)|, a bin-weighted average gap between accuracy and confidence. The senior caveat: ECE is binning-dependent (10–15 bins is standard) — quoting an absolute ECE without the binning scheme is meaningless, and variants like ACE/TACE use adaptive bins. The publishable bar for top-confidence buckets is roughly ECE < 0.01.
ECE’s sibling is the Brier score — the mean squared error between the predicted probability and the 0/1 outcome. Unlike ECE it is a proper scoring rule: it is minimized only by the true probabilities and rewards both calibration and sharpness (confident-and-correct), so it cannot be gamed by always predicting the base rate. The clean decomposition (Murphy) splits the Brier score into reliability (calibration error) minus resolution plus the irreducible uncertainty. Interview angle. “ECE vs Brier?” → ECE is an intuitive diagnostic but binning-dependent and not strictly proper; the Brier score (or log loss) is the proper scoring rule you optimize and report when you need a single trustworthy number.
KDD 2020 Tutorial: How to Calibrate Your Neural Network ClassifierACM

When calibration matters — and when it is wasted effort

Interview angle. “Is your model calibrated, and does it matter?” — the strong answer does not calibrate by default; it asks whether the use case is probability-sensitive. Calibrate when a threshold moves (review budgets, cost-sensitive cutoffs), when you stack/blend models with mismatched score distributions, or when you communicate risk to a human. Skip it when the operating point is fixed and the KPI is pure ranking (NDCG/MAP) or AUC. Distinguishing “I need calibration because the threshold will move” from “I don’t, I’m locking precision@500” is exactly the signal Dan Lee’s fraud prompt rewards.
The mechanics matter too: calibration must be fit on a held-out fold the model never trained on (scikit-learn’s CalibratedClassifierCV does this with an internal CV split), because calibrating on the training data just re-learns the same over-confidence. And calibration is monotone — it re-maps scores without reordering them — so it changes Brier/ECE but leaves AUC and the ranking essentially fixed. That is the whole reason it is a safe, cheap post-processing step: you can always add it to make probabilities honest without risking your discrimination metric.

Interview prep

This lesson hands you two adjacent question clusters. For regularization, lead with geometry and pick a side by feature statistics. For calibration, lead with the reliability diagram and the orthogonality to AUC, and treat calibration as conditional. Have these ready.
  1. 01“L1 vs L2?” → L1 (diamond) zeros weights → sparsity/selection; L2 (sphere) shrinks all → handles correlated/ill-conditioned features.
  2. 02“What if features are correlated?” → lasso keeps one arbitrarily; ridge shrinks the group together; elastic-net selects-and-shrinks.
  3. 03“Why L2 (weight decay) in nets, not L1?” → L1’s non-smoothness fights SGD; L2 integrates cleanly with gradient flow.
  4. 04“Early stopping — is it real regularization?” → yes, spectral: halting iteration constrains complexity; it’s the free regularizer on any val-fold run.
  5. 05“Is calibration the same as accuracy/AUC?” → no — it’s about probability quality and is orthogonal to ranking; calibrating leaves AUC unchanged.
  6. 06“Platt vs isotonic?” → Platt for sigmoid distortion and < 1000 cal points; isotonic for ≥ 1000 (any monotone distortion, but overfits when small).
  7. 07“How do you measure miscalibration?” → reliability diagram + ECE = bin-weighted |accuracy − confidence|; report the binning scheme.
  8. 08“When does calibration actually matter?” → moving thresholds, cost-sensitive decisions, model blending, risk communication — not fixed-threshold ranking.
Follow-ups: “your gradient-boosted model has great AUC but the predicted probabilities cluster around 0.3–0.7 — what’s happening and how do you fix it?” → boosted trees push mass away from 0/1 (sigmoid distortion); fit Platt on a held-out fold (or isotonic if ≥ 1000 points). “Deep net is over-confident — fix?” → temperature scaling on the logits, which preserves the argmax. And the trap: “just always use isotonic, it’s more flexible” — wrong below ~1000 calibration points, where it overfits and Platt wins. Always say you would measure ECE before and after, on a fold the model never trained on.
paperPredicting Good Probabilities With Supervised Learning (Platt vs isotonic)Niculescu-Mizil & Caruana (ICML 2005)paperOn Calibration of Modern Neural Networks (temperature scaling, ECE)Guo et al. (ICML 2017)paperDropout: A Simple Way to Prevent Neural Networks from OverfittingSrivastava et al. (JMLR 2014)

Checkpoint

You have ~2,000 features, strongly suspect only a few dozen carry signal, and the business wants a short, explainable feature list. Which regularizer and why?

ARidge (L2), because it is the default and stabilizes trainingBNo regularization, then prune by p-valueCLasso (L1), because its diamond geometry drives most weights to exactly zero, yielding feature selection and a short explainable model
Sign up free to answer and see why

Checkpoint

A gradient-boosted fraud model has 0.94 ROC-AUC, but its predicted probabilities are needed to set a cost-based threshold that will shift weekly. A reliability diagram shows bars well below the diagonal. Best step?

ACalibrate the scores (Platt if < 1000 calibration points, else isotonic) on a held-out fold, and verify ECE drops — AUC is fine, the probabilities are notBRetrain with a larger model to raise AUCCLower the decision threshold to 0.3 to compensate for the over-confidence
Sign up free to answer and see why

Checkpoint

A model produces a group of three highly correlated features. Lasso keeps one and zeros the other two, and the chosen one flips between resamples. What is the most stable fix that still controls variance?

AIncrease the lasso λ until only one feature survives everywhereBDrop two of the three features by hand before fittingCUse elastic-net (or ridge) so the correlated group is shrunk together rather than split arbitrarily
Sign up free to answer and see why

Checkpoint

You have only ~400 examples in your held-out calibration set and need to calibrate a classifier whose distortion looks sigmoidal. Which method, and what’s the risk of the alternative?

AIsotonic regression, because it is strictly more flexible than PlattBPlatt scaling, because for < 1000 calibration points it beats isotonic, which would overfit; the sigmoidal distortion also matches Platt’s formCNo calibration — 400 points is too few for either
Sign up free to answer and see why

Checkpoint

A teammate adds L2, dropout, AND aggressive early stopping to a model that was only mildly overfitting; now it underfits. What’s the principled correction?

AKeep all three but train for many more epochs to compensateBSwitch entirely to a bigger model so the extra regularization is absorbedCRecognize each regularizer adds bias; for a mild overfit, dial back to one well-chosen regularizer and tune its strength to the smallest that closes the gap
Sign up free to answer and see why

Could you pick L1/L2/elastic-net from feature statistics, explain why dropout/early-stopping regularize, and decide Platt-vs-isotonic and whether calibration even matters?

New to itGetting thereConfident

Takeaways

  • Regularization is the variance dial: L1 (diamond) → sparsity/selection; L2 (sphere) → smooth shrinkage for correlated features; elastic-net mixes both.
  • Dropout ≈ averaging an exponential ensemble (MNIST 1.6→1.35%); early stopping is spectral regularization and free on any val fold.
  • Match one regularizer to the prior; stacking them blindly just adds bias.
  • Calibration is orthogonal to AUC: it fixes probability quality (ECE/Brier), not ranking.
  • Platt for < 1000 calibration points / sigmoid distortion; isotonic for ≥ 1000; temperature for deep-net logits — and only calibrate when a decision consumes the probability.

Next: model selection in practice — logistic vs gradient-boosted trees vs neural nets vs LLMs, and when to use which.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.