Lesson 6 of 6 · 45 min
Mock modeling interview
A timed rapid-fire gauntlet across the whole track — bias-variance, loss/optimization, regularization, calibration, model selection, metrics, leakage. Practitioner framing (process over answer), the khangich/alirezadir question banks, and five scenario checkpoints under pressure.
How the modeling round is actually scored
Round 0 — the answer template that scores on every question
1THE UNIVERSAL ANSWER TEMPLATE (say it in this order)23 1. AXIS "This is a question."4 2. CONSTRAINT "The binding constraint is ."5 3. DIAGNOSIS "I'd first check ."6 4. LEVER "So I'd turn in , because ."7 5. VERIFY "And I'd confirm it helped by ."89 Definition-only answers skip 2-5. The hire signal is steps 4 and 5.
Machine Learning Fundamentals: Bias and Variance (warm-up recap)StatQuest with Josh StarmerRound 1 — rapid-fire fundamentals (60–90s each)
- 01“Bias-variance in one breath?” → expected error = noise + bias² + variance; more data kills variance not bias; diagnose from the train/val gap.
- 02“Overfitting — detect and fix?” → train ≪ val gap; fix with regularization, more data, simpler model, or early stopping — name the lever.
- 03“Cross-entropy vs MSE for classification?” → CE is the label-distribution NLL; MSE+sigmoid vanishes gradients on confident errors and is non-convex there.
- 04“L1 vs L2?” → L1 zeros weights (sparsity); L2 shrinks all (correlated/ill-conditioned); say which by feature statistics.
- 05“SGD vs Adam — when not Adam?” → Adam adapts per-coordinate (transformers); SGD+momentum for SOTA vision (Adam can generalize worse).
- 06“Is your model calibrated, does it matter?” → reliability diagram + ECE; only calibrate when a threshold moves / costs are consumed — orthogonal to AUC.
- 07“ROC-AUC vs PR-AUC?” → PR-AUC / Precision@K for rare positives; ROC hides precision collapse via FPR’s huge denominator.
- 08“Most dangerous leakage?” → temporal/group; fix with forward-chaining + GroupKFold and features observable at prediction time.
Key idea
Round 2 — the scenario probes (where points are won)
- 01“CART overfits 10M rows — first knob?” → cap max_depth / raise min_samples_leaf (bias↑, variance↓); rule out leakage via group/time CV; verify on the val curve.
- 02“GBT train-AUC 0.93 / val 0.78, gap worst on rare campaign_id — fix?” → regularize THOSE features (min_child_weight↑, hashing, per-feature penalty), not the global model; check recall on a known slice.
- 03“0.2% fraud, 500 reviews/day — metric + threshold?” → Precision@500 / Recall@500; threshold to yield 500 alerts on recent validation; calibrate if the cutoff moves.
- 04“Amazon product-return with delayed labels — target + loss?” → label = return within a fixed post-delivery window; train only on matured labels; weighted/cost-sensitive loss; calibrate.
- 05“Next-day churn split with logs + tickets + marketing — avoid leakage?” → define prediction time t and label window; split by time; group by user; audit post-outcome features.
- 06“Offline PR-AUC improved but online engagement dropped — debug?” → verify eval parity (dataset version, label def, joins); backtest old vs new through each pipeline; check drift/label-delay; bootstrap CIs.
- 07“Spam model degrades in production — fix?” → drop short-lived n-grams, hash rare tokens, raise L2; validate with rolling-window CV that matches production.
- 08“Baseline for a new churn model?” → calibrated logistic / stratified prior + a tenure-and-activity rule; it sets a strong bar and surfaces label/feature bugs.
When two valid tools both fit, the interviewer is not testing whether you know them — they already assume you do. They are testing whether you can say which one you’d reach for first, why, and how you’d know if you were wrong. Pick a side and bring a verification step.
Round 3 — the pen-and-paper / derivation asks
(ŷ−y)x, argue the logistic Hessian XᵀWX is PSD (hence convex), and explain why MSE saturates against a sigmoid. Chip Huyen’s counter-rubric — “prioritize intuitions over equations” — is not a contradiction: the winning candidate can both derive it and explain it in words, with the explanation the deciding factor. So practice each derivation until you can narrate it, not just write it.1THE FOUR DERIVATIONS TO HAVE COLD23 1. Bias-variance: E[(y - f_hat)^2] = sigma^2 + bias^2 + variance4 cross-terms vanish: f_hat _|_ epsilon (E=0), and E[E f_hat - f_hat]=0.56 2. Logistic gradient: d/dtheta CE = (sigmoid(theta.x) - y) * x (error * input).78 3. Convexity: Hessian = X^T W X, W=diag[p(1-p)] >= 0 => PSD => convex;9 full-rank X => unique minimum.1011 4. MSE-vs-sigmoid saturation: MSE grad carries sigma'(z)=p(1-p) ~ 0 on12 confident mistakes; CE cancels it, keeping grad ~ (p - y).Round 4 — the “now productionize it” twist
The modeling round is won by the candidate who, asked “what would you do,” answers with a specific lever and a way to check it — and who, asked “and then in production,” already has the monitoring and the fallback ready. Process over answer, all the way down.
Interview prep
- 01“Walk me through diagnosing an overfit model.” → train/val gap → name the lever and its bias/variance direction → re-check the curve.
- 02“Why this loss?” → state the noise model (Bernoulli/Gaussian) → what the gradient does when wrong → the production implication.
- 03“Pick a regularizer for these features.” → geometry + feature statistics (sparse→L1, correlated→L2/elastic-net) → smallest λ that closes the gap.
- 04“Does this use case need calibration?” → only if a threshold moves / costs are consumed / models are blended; measure ECE before and after.
- 05“Choose a model for this problem.” → data shape + latency first → calibrated baseline → climb only on measured lift past the maintenance cost.
- 06“Which metric, given this constraint?” → name the business limit → the metric that lives at it (Precision@K for capacity, PR-AUC for rare positives).
- 07“Set the threshold.” → calibrate → cut at 0.5 (symmetric) or p* = C(FP)/(C(FP)+C(FN)) (asymmetric), not Youden’s J.
- 08“This model looks too good — debug.” → trace feature timestamps, fit preprocessing on train only, GroupKFold/forward-chain, remove top feature to expose leakage.
Common mistake
The #1 red-flag pattern across the whole round: answering with a definition and stopping.
Checkpoint
Rapid-fire: “Your model overfits — what do you do?” Under time pressure, which answer best matches the hiring rubric?
Checkpoint
Scenario: “0.2% fraud, ops reviews 500 alerts/day.” The interviewer asks for your metric and threshold. Strongest response?
Checkpoint
Derivation ask: “Show me why logistic loss is convex.” With the whiteboard marker in hand, the best move is:
Checkpoint
Scenario: “Offline PR-AUC went up after your change, but the online A/B test shows engagement dropped.” Best debugging path?
Checkpoint
Scenario: “Predict next-day churn from event logs, support tickets, and marketing touches — design the split to avoid leakage while using as much data as possible.” Strongest answer?
Could you run a full modeling round — rapid-fire in 90s each, scenarios as axis→constraint→lever→verify, and the four derivations cold?
Takeaways
- The round scores process over answer: open every question by naming the axis, the constraint, and the verification step.
- Every answer has the shape mechanism → lever → verification; memorize the shape, not just the facts.
- Scenarios reward sequencing and diagnosis; when a popular default exists (SMOTE, “calibrate everything,” ROC-AUC), say why it’s not your first move.
- Be able to BOTH derive (bias-variance cancellation, logistic gradient/Hessian, MSE saturation) and explain it in words.
- The universal red flag is the definition-only answer; convert each into a decision plus how you’d check it.
You’ve completed ML Fundamentals for Interviews — revisit any lesson’s checkpoints to keep the levers sharp before the real round.
Sources
Free to read · better with Enzo
Learn it with Enzo
Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.