Lesson 1 of 6 · 48 min

Bias-variance & generalization

The decomposition with the actual math, the three terms and what drives each, why the modern interview asks it as a sizing question (“which hyperparameter first?”), and diagnosing under/overfitting from a learning curve instead of a single number.

Why this is the master equation of the modeling round

Every other lesson in this track is a knob on one equation: expected test error = irreducible noise + bias² + variance. Regularization moves variance down at a bias cost; model selection trades the two across families; calibration and metrics are how you measure where you landed. The interview almost never asks “define bias and variance” anymore — the 2026 formulation (DataInterview, Dan Lee) is operational: “your CART overfits 10M rows, which hyperparameter do you change first, and which way does it move bias and variance?” A definition-only answer is the universal killer here. This lesson gives you the decomposition cold, then the levers and the curve-reading that turn it into a hire signal.
Set up the only assumption you need: a label is generated as y = f(x) + ε with E[ε]=0 and Var(ε)=σ². You fit an estimate f̂ on a finite training set S drawn at random. The question the decomposition answers is: across all the training sets you might have drawn, what is the expected squared error of f̂(x₀) at a fixed test point x₀? That “across all training sets” framing is the whole trick — bias and variance are both expectations over the randomness of S, not properties of one fit.
The result (ESL ch. 7; CS229 notes) is exact, not a bound: Err(x₀) = σ² + [E f̂(x₀) − f(x₀)]² + E[ f̂(x₀) − E f̂(x₀) ]². Read left to right: σ² is irreducible — noise in y you cannot predict by definition; bias² is how far the average model E f̂ sits from truth f (structural misspecification); variance is how much a single fit bounces around that average as S changes. The cross-terms vanish because f̂ is independent of ε and E[E f̂ − f̂]=0 — that cancellation is the derivation, and being able to say it out loud is the senior tell.
Make it concrete with the two canonical knobs. Fit a polynomial of degree d to a wiggly truth: degree 1 underfits (high bias, the line cannot bend to the curve, but it barely moves between samples — low variance); degree 15 interpolates every point (near-zero bias on average, but each new sample yanks the curve wildly — high variance). Same story for k-NN: k=1 is maximum variance (the prediction is one noisy neighbour), k=N collapses to the global mean (maximum bias, zero variance). The whole field of model selection is just choosing d or k — the complexity — to sit at the bottom of the summed-error U-curve.
Machine Learning Lecture 19: Bias-Variance Decomposition (Cornell CS4780)Kilian Weinberger

The three terms, and the one lever each responds to

CS229 pins each term to a mechanism, and this mapping is what an interviewer wants back. Irreducible noise is “unavoidable — we cannot predict ε by definition”; no model change moves it, only richer features x that explain away what looked like noise. Bias² “captures the part of the error introduced by the lack of expressivity of the model … not due to lack of data, but because the family fundamentally cannot approximate” the target — so you attack it with capacity (deeper trees, more features, a richer family). Variance “captures how the random nature of the finite dataset introduces errors” and “often decreases as the dataset size increases” — so you attack it with more data, regularization, bagging, or early stopping.
Bias is what the model gets wrong on average; variance is how much it changes its mind when you change the data. You cannot drive both to zero with a fixed budget — you can only choose where to spend it. That sentence is the entire modeling round in one line.
code
1THE DECOMPOSITION, AS A LEVER TABLE23  term            driven by                         the knob that shrinks it4  -------------   ------------------------------     --------------------------------5  sigma^2         label noise, missing features      better features (NOT model class)6  bias^2          family too rigid for true f        more capacity / richer family7  variance        capacity high OR sample small      more data, regularize, bag, early-stop89  Increasing complexity:  bias^2 DOWN, variance UP.  Decreasing complexity: the reverse.10  Cross-validation is just the procedure that finds where their sum is minimized.
The irreducible term has a name worth dropping: in classification it is the Bayes error — the error of the optimal classifier that knows the true p(y|x) exactly. No model, however large, beats it; it is the floor set by genuine label ambiguity and missing features. Interview angle. “Can you get to 0% error?” → only if Bayes error is 0 (deterministic labels, fully informative features). If two inputs with identical features have different labels in your data, you have nonzero irreducible error and chasing it with capacity just buys variance. Recognizing an irreducible floor — rather than blaming the model — is a senior reflex.
Interview angle. The intermediate version of the question is “why does adding complexity reduce bias but raise variance?” Answer with the capacity argument: a richer family contains more functions, so the best-fitting one sits closer to f (less bias) — but it also has more freedom to chase the noise in this particular S, so it bounces more across draws (more variance). Then name the concrete artifact: a tree’s max_depth, a forest’s max_features, a polynomial’s degree, a network’s width. Naming the artifact is what separates the strong answer from the textbook one.

Diagnosing from the curve, not from a number

The accepted detection method (Data Science Stack Exchange; Yuan Meng’s MLE writeup) is a gap, never a single score. Compare training error with validation/CV error and read the shape: both high → underfit (high bias — the family is too weak, or you over-regularized); train low, validation high → overfit (high variance — the model memorized this sample). A weak candidate quotes “low training error” as if it were good news; a strong one says “low training error with a large val gap is the signature of overfitting, and low training error with a small gap is the goal.”
code
1LEARNING-CURVE READING (train vs validation error)23  both high, close together      ->  UNDERFIT (bias)     -> add capacity / features4  train low, val high, big gap   ->  OVERFIT (variance)  -> regularize / more data / simpler5  both low, small gap            ->  about right         -> ship; tune at the margin6  val U-turns up while train      ->  overfitting in TIME -> early stop at the val minimum7    keeps falling89  Rule: more data flattens variance but CANNOT fix bias -- a too-simple10  family stays wrong no matter how many rows you feed it.
The “more data” nuance is a favorite follow-up: “you are overfitting — will more data fix it?” The honest answer is it depends which term dominates. More data drives variance toward zero (the fit stops depending on the particular sample), but it does nothing to bias — a linear model on a quadratic truth is equally wrong at 1k or 1M rows. So if your learning curves have already converged and both sit high, you are bias-bound and need a richer model, not more rows. Interview angle. Sketching the two learning curves and pointing at the gap (or its absence) is the single most convincing way to answer any over/underfitting question.

The sizing question: which knob first, and which way

The canonical 2026 prompt (Dan Lee): “you trained a CART on 10M rows and it overfits badly — what do you change first, and how does each change affect bias and variance?” The model answer names levers and directions: cap growth with max_depth, raise min_samples_leaf and min_samples_split (each forces more averaging per leaf → more bias, less variance); tune cost-complexity pruning ccp_alpha or min_impurity_decrease; and — the senior beat — “validate with time- or group-aware CV, because the apparent overfit might be leakage rather than pure capacity.” That last clause ties Lesson 1 to Lesson 5 and signals you have shipped models, not just read about them.
The ensemble follow-up tests whether you know how the two families move the tradeoff in opposite ways. Bagging / random forests average many high-variance, low-bias trees → variance falls (decorrelation via max_features is the lever), bias roughly unchanged. Boosting (gradient-boosted trees) fits shallow, high-bias learners sequentially to the residual → bias falls, but it can raise variance and overfit if you add too many rounds, which is why early stopping on a holdout and a small learning_rate are standard. So “random forest overfits → cut tree depth and max_features; GBT overfits → cut n_estimators / learning rate and early-stop” is the crisp, lever-level answer.

Estimating the two terms — and reading the regularization path

Since bias and variance are expectations over training sets, you estimate them empirically: repeatedly resample the training data (bootstrap or repeated CV), refit, and at each test point measure how far the average prediction sits from the target (bias²) and how much individual predictions scatter around that average (variance). Libraries like mlxtend’s bias_variance_decomp automate exactly this. The practical read: if the scatter across resamples is wide, you are variance-bound; if every resample is confidently wrong in the same direction, you are bias-bound. This is the rigorous version of “sketch the two learning curves.”
The same lens explains the regularization path you will be asked to read. Sweep a ridge penalty λ from 0 upward and watch validation error trace a U: at λ=0 the model is unconstrained (low bias, high variance — the right arm of the U); as λ grows it shrinks coefficients (variance falls, bias rises) until, as CS229 warns, “an extremely large λ can result in a model with large bias” (the left arm). The minimum of that U is the bias-variance sweet spot — and it is literally what cross-validated grid search over λ finds. Interview angle. “What does the validation curve over λ tell you?” → the descending arm is variance you are buying back as bias; the ascending arm is bias you have over-bought; ship the λ at the trough.

Where the decomposition stops being clean

A precise candidate knows the limits of the equation. The additive bias²+variance split is specific to squared loss on a real-valued target. Under 0/1 classification loss it does not decompose additively — the influential Domingos (2000) treatment shows variance can be helpful on the wrong side of the decision boundary, so “high variance is always bad” is false for classifiers. That is also why for classification we reason about generalization gap, calibration, and the metrics of Lesson 5 rather than a literal variance number. Mentioning this unprompted reads as genuine depth.
One more framing that lands well: bias-variance is the statistical face of the No-Free-Lunch theorem. No single capacity wins on every problem — the right point on the curve is dataset-dependent, which is exactly why Lesson 4’s case studies (a seasonal random walk beating XGBoost on small EPS data; XGBoost crushing deep nets on tabular) are predicted by this lesson rather than surprising. The decomposition tells you what to trade; the data tells you where.

Interview prep

Bias-variance is a guaranteed opener for a data-scientist modeling round. Lead with the mechanism, name a concrete lever and its direction, and finish with how you would verify the change on a learning curve. Have these eight crisp answers ready.
  1. 01“State the decomposition.” → expected test error = irreducible noise σ² + bias² + variance; bias² and variance are both expectations over the random training set.
  2. 02“Which term does more data fix?” → variance (it converges as n grows); bias is structural and unmoved by more rows.
  3. 03“CART overfits 10M rows — first knob?” → lower max_depth / raise min_samples_leaf (more bias, less variance); then check it is not leakage via group/time-aware CV.
  4. 04“Why does complexity cut bias but raise variance?” → richer family fits f closer on average, but also chases this sample’s noise more → bounces across draws.
  5. 05“Under- vs overfit from a curve?” → both errors high = underfit/bias; train low + val high (big gap) = overfit/variance.
  6. 06“Bagging vs boosting on the tradeoff?” → bagging averages to cut variance; boosting fits residuals to cut bias (and can overfit, so early-stop).
  7. 07“Does the decomposition hold for classification?” → not additively under 0/1 loss; variance can even help past the boundary (Domingos 2000).
  8. 08“Bigger nets generalize better — broken tradeoff?” → no; double descent, with implicit regularization (SGD/weight decay) controlling effective capacity.
Common follow-ups chain fast: from “which knob” to “show me the bias/variance direction,” then “how would you confirm it helped” (re-plot the gap; watch validation, not training), then the leakage curveball (“the gap is worst on a rare categorical ID like campaign_id — now what?”). For that last one the strong move is feature-targeted, not global: regularize the rare IDs (higher min_child_weight, hashing, per-feature penalty) rather than reaching for a smaller global model. Always close on verification — interviewers (Exponent, Dan Lee) grade “process over answer.”
paperThe Elements of Statistical Learning, ch. 7 (Model Assessment & Selection)Hastie, Tibshirani & FriedmanvideoLecture 08 — Bias-Variance Tradeoff (Learning From Data)Yaser Abu-Mostafa, CaltechpaperA Unified Bias-Variance Decomposition (incl. 0/1 loss)Pedro Domingos

Checkpoint

Your random forest gets 0.99 train AUC and 0.78 validation AUC; the learning curves have clearly plateaued and the gap is stable. An interviewer asks what you change first. Best move?

AReduce variance: lower max_depth and max_features, and add trees until validation stabilizes — the large stable gap is overfitting, not biasBSwitch to a higher-capacity model like a deep net to close the gapCCollect 10× more data, since more data always fixes a train/validation gap
Sign up free to answer and see why

Checkpoint

A linear model gives 0.42 RMSE on train and 0.43 on validation. Adding 5× more data barely moves either. What is going on and what helps?

AOverfitting — apply L2 regularization to shrink the gapBUnderfitting (high bias) — the family is too rigid; add capacity or richer features, since more data cannot fix biasCThe validation set is too small — increase its size to reveal the real gap
Sign up free to answer and see why

Checkpoint

An interviewer says: “Prove you understand the decomposition — why do the cross-terms vanish?” Strongest answer?

AThey cancel because bias and variance are defined to be orthogonal by conventionBBecause squared loss is convex, so cross-terms are always zeroCf̂ is independent of the noise ε (so E[(f−f̂)ε]=0), and E[E f̂ − f̂]=0 makes the bias×deviation cross-term vanish — leaving σ² + bias² + variance
Sign up free to answer and see why

Checkpoint

You train a gradient-boosted tree; validation loss falls, bottoms out at round 380, then rises while training loss keeps dropping. Best interpretation and action?

AThe learning rate is too low — raise it so training finishes fasterBOverfitting in time as rounds accumulate — early-stop at the validation minimum (~round 380) and/or lower the learning rateCIrreducible noise has been reached — nothing more can be done
Sign up free to answer and see why

Checkpoint

A teammate claims “our deep net beats the logistic baseline because over-parameterized models just don’t overfit.” How should you push back in an interview-credible way?

AAgree — with enough parameters the bias-variance tradeoff no longer appliesBDisagree entirely — over-parameterized nets always overfit and should be avoidedCNuance it: the tradeoff still holds, but double descent plus implicit regularization (SGD, weight decay, early stopping) controls effective capacity — the win is from low-norm solutions, not from “not overfitting”
Sign up free to answer and see why

Could you state the decomposition cold, name the lever (and its bias/variance direction) for an over/underfit model, and read the diagnosis off a learning curve?

New to itGetting thereConfident

Takeaways

  • Expected error = irreducible noise σ² + bias² + variance; bias and variance are expectations over the random training set, not properties of one fit.
  • More data kills variance, never bias; bias needs a richer family or better features.
  • Diagnose from the train/validation gap, not a single number: both high = bias, big gap = variance.
  • Answer the sizing question with levers AND directions (max_depth↑ → bias↑/variance↓), then verify on the curve.
  • 0/1 loss does not decompose additively; over-parameterized nets show double descent under implicit regularization — neither repeals the tradeoff.

Next: loss functions & optimization — why cross-entropy not MSE for classification, convexity, and SGD vs Adam.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.