Lesson 1 of 6 · 48 min
Bias-variance & generalization
The decomposition with the actual math, the three terms and what drives each, why the modern interview asks it as a sizing question (“which hyperparameter first?”), and diagnosing under/overfitting from a learning curve instead of a single number.
Why this is the master equation of the modeling round
y = f(x) + ε with E[ε]=0 and Var(ε)=σ². You fit an estimate f̂ on a finite training set S drawn at random. The question the decomposition answers is: across all the training sets you might have drawn, what is the expected squared error of f̂(x₀) at a fixed test point x₀? That “across all training sets” framing is the whole trick — bias and variance are both expectations over the randomness of S, not properties of one fit.Err(x₀) = σ² + [E f̂(x₀) − f(x₀)]² + E[ f̂(x₀) − E f̂(x₀) ]². Read left to right: σ² is irreducible — noise in y you cannot predict by definition; bias² is how far the average model E f̂ sits from truth f (structural misspecification); variance is how much a single fit bounces around that average as S changes. The cross-terms vanish because f̂ is independent of ε and E[E f̂ − f̂]=0 — that cancellation is the derivation, and being able to say it out loud is the senior tell.d to a wiggly truth: degree 1 underfits (high bias, the line cannot bend to the curve, but it barely moves between samples — low variance); degree 15 interpolates every point (near-zero bias on average, but each new sample yanks the curve wildly — high variance). Same story for k-NN: k=1 is maximum variance (the prediction is one noisy neighbour), k=N collapses to the global mean (maximum bias, zero variance). The whole field of model selection is just choosing d or k — the complexity — to sit at the bottom of the summed-error U-curve.
Machine Learning Lecture 19: Bias-Variance Decomposition (Cornell CS4780)Kilian WeinbergerThe three terms, and the one lever each responds to
x that explain away what looked like noise. Bias² “captures the part of the error introduced by the lack of expressivity of the model … not due to lack of data, but because the family fundamentally cannot approximate” the target — so you attack it with capacity (deeper trees, more features, a richer family). Variance “captures how the random nature of the finite dataset introduces errors” and “often decreases as the dataset size increases” — so you attack it with more data, regularization, bagging, or early stopping.Bias is what the model gets wrong on average; variance is how much it changes its mind when you change the data. You cannot drive both to zero with a fixed budget — you can only choose where to spend it. That sentence is the entire modeling round in one line.
1THE DECOMPOSITION, AS A LEVER TABLE23 term driven by the knob that shrinks it4 ------------- ------------------------------ --------------------------------5 sigma^2 label noise, missing features better features (NOT model class)6 bias^2 family too rigid for true f more capacity / richer family7 variance capacity high OR sample small more data, regularize, bag, early-stop89 Increasing complexity: bias^2 DOWN, variance UP. Decreasing complexity: the reverse.10 Cross-validation is just the procedure that finds where their sum is minimized.p(y|x) exactly. No model, however large, beats it; it is the floor set by genuine label ambiguity and missing features. Interview angle. “Can you get to 0% error?” → only if Bayes error is 0 (deterministic labels, fully informative features). If two inputs with identical features have different labels in your data, you have nonzero irreducible error and chasing it with capacity just buys variance. Recognizing an irreducible floor — rather than blaming the model — is a senior reflex.f (less bias) — but it also has more freedom to chase the noise in this particular S, so it bounces more across draws (more variance). Then name the concrete artifact: a tree’s max_depth, a forest’s max_features, a polynomial’s degree, a network’s width. Naming the artifact is what separates the strong answer from the textbook one.Key idea
Diagnosing from the curve, not from a number
1LEARNING-CURVE READING (train vs validation error)23 both high, close together -> UNDERFIT (bias) -> add capacity / features4 train low, val high, big gap -> OVERFIT (variance) -> regularize / more data / simpler5 both low, small gap -> about right -> ship; tune at the margin6 val U-turns up while train -> overfitting in TIME -> early stop at the val minimum7 keeps falling89 Rule: more data flattens variance but CANNOT fix bias -- a too-simple10 family stays wrong no matter how many rows you feed it.The sizing question: which knob first, and which way
max_depth, raise min_samples_leaf and min_samples_split (each forces more averaging per leaf → more bias, less variance); tune cost-complexity pruning ccp_alpha or min_impurity_decrease; and — the senior beat — “validate with time- or group-aware CV, because the apparent overfit might be leakage rather than pure capacity.” That last clause ties Lesson 1 to Lesson 5 and signals you have shipped models, not just read about them.max_features is the lever), bias roughly unchanged. Boosting (gradient-boosted trees) fits shallow, high-bias learners sequentially to the residual → bias falls, but it can raise variance and overfit if you add too many rounds, which is why early stopping on a holdout and a small learning_rate are standard. So “random forest overfits → cut tree depth and max_features; GBT overfits → cut n_estimators / learning rate and early-stop” is the crisp, lever-level answer.Estimating the two terms — and reading the regularization path
mlxtend’s bias_variance_decomp automate exactly this. The practical read: if the scatter across resamples is wide, you are variance-bound; if every resample is confidently wrong in the same direction, you are bias-bound. This is the rigorous version of “sketch the two learning curves.”λ from 0 upward and watch validation error trace a U: at λ=0 the model is unconstrained (low bias, high variance — the right arm of the U); as λ grows it shrinks coefficients (variance falls, bias rises) until, as CS229 warns, “an extremely large λ can result in a model with large bias” (the left arm). The minimum of that U is the bias-variance sweet spot — and it is literally what cross-validated grid search over λ finds. Interview angle. “What does the validation curve over λ tell you?” → the descending arm is variance you are buying back as bias; the ascending arm is bias you have over-bought; ship the λ at the trough.Common mistake
“Deep learning broke the bias-variance tradeoff — bigger models just generalize better.”
Where the decomposition stops being clean
Key idea
Interview prep
- 01“State the decomposition.” → expected test error = irreducible noise σ² + bias² + variance; bias² and variance are both expectations over the random training set.
- 02“Which term does more data fix?” → variance (it converges as n grows); bias is structural and unmoved by more rows.
- 03“CART overfits 10M rows — first knob?” → lower max_depth / raise min_samples_leaf (more bias, less variance); then check it is not leakage via group/time-aware CV.
- 04“Why does complexity cut bias but raise variance?” → richer family fits f closer on average, but also chases this sample’s noise more → bounces across draws.
- 05“Under- vs overfit from a curve?” → both errors high = underfit/bias; train low + val high (big gap) = overfit/variance.
- 06“Bagging vs boosting on the tradeoff?” → bagging averages to cut variance; boosting fits residuals to cut bias (and can overfit, so early-stop).
- 07“Does the decomposition hold for classification?” → not additively under 0/1 loss; variance can even help past the boundary (Domingos 2000).
- 08“Bigger nets generalize better — broken tradeoff?” → no; double descent, with implicit regularization (SGD/weight decay) controlling effective capacity.
campaign_id — now what?”). For that last one the strong move is feature-targeted, not global: regularize the rare IDs (higher min_child_weight, hashing, per-feature penalty) rather than reaching for a smaller global model. Always close on verification — interviewers (Exponent, Dan Lee) grade “process over answer.”Common mistake
The #1 red-flag answer: “Bias is error from wrong assumptions, variance is sensitivity to data, and there’s a tradeoff.” (…and then stopping.)
min_samples_leaf, which adds bias and cuts variance, and I’d confirm on the validation curve.”Checkpoint
Your random forest gets 0.99 train AUC and 0.78 validation AUC; the learning curves have clearly plateaued and the gap is stable. An interviewer asks what you change first. Best move?
Checkpoint
A linear model gives 0.42 RMSE on train and 0.43 on validation. Adding 5× more data barely moves either. What is going on and what helps?
Checkpoint
An interviewer says: “Prove you understand the decomposition — why do the cross-terms vanish?” Strongest answer?
Checkpoint
You train a gradient-boosted tree; validation loss falls, bottoms out at round 380, then rises while training loss keeps dropping. Best interpretation and action?
Checkpoint
A teammate claims “our deep net beats the logistic baseline because over-parameterized models just don’t overfit.” How should you push back in an interview-credible way?
Could you state the decomposition cold, name the lever (and its bias/variance direction) for an over/underfit model, and read the diagnosis off a learning curve?
Takeaways
- Expected error = irreducible noise σ² + bias² + variance; bias and variance are expectations over the random training set, not properties of one fit.
- More data kills variance, never bias; bias needs a richer family or better features.
- Diagnose from the train/validation gap, not a single number: both high = bias, big gap = variance.
- Answer the sizing question with levers AND directions (max_depth↑ → bias↑/variance↓), then verify on the curve.
- 0/1 loss does not decompose additively; over-parameterized nets show double descent under implicit regularization — neither repeals the tradeoff.
Next: loss functions & optimization — why cross-entropy not MSE for classification, convexity, and SGD vs Adam.
Sources
Free to read · better with Enzo
Learn it with Enzo
Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.