Lesson 6 of 6 · 45 min

Mock modeling interview

A timed rapid-fire gauntlet across the whole track — bias-variance, loss/optimization, regularization, calibration, model selection, metrics, leakage. Practitioner framing (process over answer), the khangich/alirezadir question banks, and five scenario checkpoints under pressure.

How the modeling round is actually scored

A real ML theory round (Priyaa’s FAANG writeup: 45–60 minutes; Stripe’s loop: recruiter → OA → ML bug-squash → ML design → HM) is not a definitions quiz — it’s a prioritization test. Chip Huyen states the load-bearing rubric outright: “interviewers often care more about your approach than the actual objective correctness of your answers,” and “strong answers should prioritize intuitive explanations over mathematical equations.” Every topic in this track has the same weak-answer pattern (define the term and stop) and the same strong-answer pattern (name the axis → sequence the levers → attach a verification step). This lesson drills the muscle memory under time.

Round 0 — the answer template that scores on every question

The meta-frame to open with (spend the first ~90 seconds here): name (a) the axis you’re operating on — loss, metric, capacity, calibration, time/leakage; (b) the constraint — business capacity, latency, regulatory, deployment distribution; and (c) the verification step you’ll run. Dan Lee’s Meta-flavored rubric is explicit: “avoid vague statements like ‘use a more complex model’; instead describe specific levers like regularization, thresholding, calibration, or better negative sampling.” That sentence is the difference between a hire and a no-hire, and it applies to every question below.
code
1THE UNIVERSAL ANSWER TEMPLATE (say it in this order)23  1. AXIS        "This is a  question."4  2. CONSTRAINT  "The binding constraint is ."5  3. DIAGNOSIS   "I'd first check ."6  4. LEVER       "So I'd turn  in , because ."7  5. VERIFY      "And I'd confirm it helped by ."89  Definition-only answers skip 2-5. The hire signal is steps 4 and 5.
Machine Learning Fundamentals: Bias and Variance (warm-up recap)StatQuest with Josh Starmer

Round 1 — rapid-fire fundamentals (60–90s each)

These are the warm-ups from the khangich and alirezadir banks. Answer each leading with the mechanism, then the one-line implication. If you can’t do all of these in under 90 seconds, that’s your review list before the next interview.
  1. 01“Bias-variance in one breath?” → expected error = noise + bias² + variance; more data kills variance not bias; diagnose from the train/val gap.
  2. 02“Overfitting — detect and fix?” → train ≪ val gap; fix with regularization, more data, simpler model, or early stopping — name the lever.
  3. 03“Cross-entropy vs MSE for classification?” → CE is the label-distribution NLL; MSE+sigmoid vanishes gradients on confident errors and is non-convex there.
  4. 04“L1 vs L2?” → L1 zeros weights (sparsity); L2 shrinks all (correlated/ill-conditioned); say which by feature statistics.
  5. 05“SGD vs Adam — when not Adam?” → Adam adapts per-coordinate (transformers); SGD+momentum for SOTA vision (Adam can generalize worse).
  6. 06“Is your model calibrated, does it matter?” → reliability diagram + ECE; only calibrate when a threshold moves / costs are consumed — orthogonal to AUC.
  7. 07“ROC-AUC vs PR-AUC?” → PR-AUC / Precision@K for rare positives; ROC hides precision collapse via FPR’s huge denominator.
  8. 08“Most dangerous leakage?” → temporal/group; fix with forward-chaining + GroupKFold and features observable at prediction time.

Round 2 — the scenario probes (where points are won)

The 2026 banks (DataInterview, Dan Lee) have moved almost entirely to scenarios. The pattern: a realistic setup, a “what do you do first / which lever” ask, then follow-ups that punish hand-waving. Walk through the canonical five out loud, each as axis → constraint → lever → verify.
  1. 01“CART overfits 10M rows — first knob?” → cap max_depth / raise min_samples_leaf (bias↑, variance↓); rule out leakage via group/time CV; verify on the val curve.
  2. 02“GBT train-AUC 0.93 / val 0.78, gap worst on rare campaign_id — fix?” → regularize THOSE features (min_child_weight↑, hashing, per-feature penalty), not the global model; check recall on a known slice.
  3. 03“0.2% fraud, 500 reviews/day — metric + threshold?” → Precision@500 / Recall@500; threshold to yield 500 alerts on recent validation; calibrate if the cutoff moves.
  4. 04“Amazon product-return with delayed labels — target + loss?” → label = return within a fixed post-delivery window; train only on matured labels; weighted/cost-sensitive loss; calibrate.
  5. 05“Next-day churn split with logs + tickets + marketing — avoid leakage?” → define prediction time t and label window; split by time; group by user; audit post-outcome features.
  6. 06“Offline PR-AUC improved but online engagement dropped — debug?” → verify eval parity (dataset version, label def, joins); backtest old vs new through each pipeline; check drift/label-delay; bootstrap CIs.
  7. 07“Spam model degrades in production — fix?” → drop short-lived n-grams, hash rare tokens, raise L2; validate with rolling-window CV that matches production.
  8. 08“Baseline for a new churn model?” → calibrated logistic / stratified prior + a tenure-and-activity rule; it sets a strong bar and surfaces label/feature bugs.
Interview angle. Notice none of these reward a single named technique — they reward sequencing and diagnosis. The SMOTE-vs-class-weights tension (lead with metric and threshold, not resampling), the calibration tension (only when the threshold moves), and the AUC-vs-PR-AUC tension (business metric over default technical metric) are all places where the “popular default” is the penalized answer. When two valid tools exist, say which you’d try first and why, then how you’d verify it helped.
When two valid tools both fit, the interviewer is not testing whether you know them — they already assume you do. They are testing whether you can say which one you’d reach for first, why, and how you’d know if you were wrong. Pick a side and bring a verification step.
A word on the failure modes that sink otherwise-strong candidates under time pressure. Rambling — listing every technique you know instead of committing to one and justifying it — reads as indecision, not breadth. Jumping to the tool before the constraint — “I’d use SMOTE / a bigger model / an LLM” before naming what’s binding — is the canonical junior tell. And going silent on a hard probe — when you genuinely don’t know, the move is to reason out loud about what you’d measure to find out, never to freeze. The clock rewards a crisp axis-constraint-lever-verify answer in 90 seconds over a meandering one that technically covers more.

Round 3 — the pen-and-paper / derivation asks

Priyaa explicitly flags pen-and-paper items that “make a huge difference”: derive the bias-variance cross-term cancellation, show the logistic gradient collapses to (ŷ−y)x, argue the logistic Hessian XᵀWX is PSD (hence convex), and explain why MSE saturates against a sigmoid. Chip Huyen’s counter-rubric — “prioritize intuitions over equations” — is not a contradiction: the winning candidate can both derive it and explain it in words, with the explanation the deciding factor. So practice each derivation until you can narrate it, not just write it.
code
1THE FOUR DERIVATIONS TO HAVE COLD23  1. Bias-variance: E[(y - f_hat)^2] = sigma^2 + bias^2 + variance4     cross-terms vanish: f_hat _|_ epsilon (E=0), and E[E f_hat - f_hat]=0.56  2. Logistic gradient: d/dtheta CE = (sigmoid(theta.x) - y) * x   (error * input).78  3. Convexity: Hessian = X^T W X, W=diag[p(1-p)] >= 0  =>  PSD  =>  convex;9     full-rank X => unique minimum.1011  4. MSE-vs-sigmoid saturation: MSE grad carries sigma'(z)=p(1-p) ~ 0 on12     confident mistakes; CE cancels it, keeping grad ~ (p - y).

Round 4 — the “now productionize it” twist

A strong derivation or scenario answer earns the operational follow-up, and this is where data-scientist candidates either show seniority or stall. The patterns to expect: “your offline metric improved but the A/B test didn’t — debug the gap” (verify evaluation parity, then backtest old vs new through each pipeline to isolate drift/label-delay); “the model is great today but degrades over weeks — what’s your monitoring?” (per-feature KS drift, per-segment PR-AUC, a scheduled retrain on a sliding window, and a calibrated baseline you can fall back to); and “labels arrive 30 days late — how do you train and evaluate?” (train only on matured labels, never treat unknowns as negatives, and use rolling-origin evaluation). The thread connecting all three is treating offline AUC as a stale number within weeks of any input-distribution change.
The modeling round is won by the candidate who, asked “what would you do,” answers with a specific lever and a way to check it — and who, asked “and then in production,” already has the monitoring and the fallback ready. Process over answer, all the way down.

Interview prep

This whole lesson is interview prep, so the “prep” section is your final-pass checklist: the eight questions that, if you can’t answer each in 60–90 seconds leading with mechanism, mark your review list. Then walk in and spend the first 90 seconds of every answer naming the axis, the constraint, and the verification.
  1. 01“Walk me through diagnosing an overfit model.” → train/val gap → name the lever and its bias/variance direction → re-check the curve.
  2. 02“Why this loss?” → state the noise model (Bernoulli/Gaussian) → what the gradient does when wrong → the production implication.
  3. 03“Pick a regularizer for these features.” → geometry + feature statistics (sparse→L1, correlated→L2/elastic-net) → smallest λ that closes the gap.
  4. 04“Does this use case need calibration?” → only if a threshold moves / costs are consumed / models are blended; measure ECE before and after.
  5. 05“Choose a model for this problem.” → data shape + latency first → calibrated baseline → climb only on measured lift past the maintenance cost.
  6. 06“Which metric, given this constraint?” → name the business limit → the metric that lives at it (Precision@K for capacity, PR-AUC for rare positives).
  7. 07“Set the threshold.” → calibrate → cut at 0.5 (symmetric) or p* = C(FP)/(C(FP)+C(FN)) (asymmetric), not Youden’s J.
  8. 08“This model looks too good — debug.” → trace feature timestamps, fit preprocessing on train only, GroupKFold/forward-chain, remove top feature to expose leakage.
Going deeper: the round is adaptive — a strong rapid-fire answer earns a harder scenario, and a strong scenario earns a derivation or a “now productionize it” twist. Treat every follow-up as an invitation to show the next layer (mechanism → scenario → derivation → ops), and when you genuinely don’t know, say what you’d measure to find out. That “here’s how I’d verify” reflex is the single highest-signal habit across every rubric in the research (Hey Amit, Exponent, Dan Lee, Stripe, Yuan Meng).
repokhangich/machine-learning-interview (FAANG question bank)Khang Phamrepoalirezadir/Machine-Learning-Interviews (ML fundamentals chapter)Alireza DirafzoondocsIntroduction to Machine Learning Interviews — “About the answers” (process over answer)Chip Huyen

Checkpoint

Rapid-fire: “Your model overfits — what do you do?” Under time pressure, which answer best matches the hiring rubric?

A“Overfitting is when the model fits the training data too closely and fails to generalize.”B“It depends — I’d confirm via the train/val gap, then turn a specific lever (e.g. raise min_samples_leaf or add L2, which adds bias and cuts variance) and re-check the validation curve to confirm it helped.”C“I’d use a more complex model so it captures the pattern better.”
Sign up free to answer and see why

Checkpoint

Scenario: “0.2% fraud, ops reviews 500 alerts/day.” The interviewer asks for your metric and threshold. Strongest response?

AOptimize ROC-AUC and pick the Youden-J threshold from the ROC curve.BMaximize overall accuracy and ship at a 0.5 threshold.CReport Precision@500 / Recall@500, set the threshold to yield 500 alerts on recent validation data, and calibrate if the cutoff must stay stable over time.
Sign up free to answer and see why

Checkpoint

Derivation ask: “Show me why logistic loss is convex.” With the whiteboard marker in hand, the best move is:

AWrite “Hessian = XᵀWX with W = diag[p(1−p)] ⪰ 0, so the loss is convex; full-rank X ⇒ unique minimum,” and narrate why p(1−p) ≥ 0 makes it PSD.BSay “it’s convex because logistic regression is a standard convex method” and move on.CPlot the sigmoid and argue that because it’s S-shaped the loss must be convex.
Sign up free to answer and see why

Checkpoint

Scenario: “Offline PR-AUC went up after your change, but the online A/B test shows engagement dropped.” Best debugging path?

AConclude the A/B test is noisy and ship the model because offline PR-AUC improved.BVerify evaluation parity (same dataset version, label definition, joins, filtering), then backtest old vs new through each pipeline to isolate where metric and behavior diverge; check drift/label-delay and bootstrap CIs.CRoll back and add more regularization until the offline and online numbers agree.
Sign up free to answer and see why

Checkpoint

Scenario: “Predict next-day churn from event logs, support tickets, and marketing touches — design the split to avoid leakage while using as much data as possible.” Strongest answer?

ARandom 80/20 split with stratification on the churn label, then k-fold CV.BDrop any feature whose name mentions churn, then random k-fold.CDefine prediction time t and the label window (churn in (t, t+1]); split by time (earlier→train, later→val/test); group by user so no user crosses the split; then audit feature importance for post-outcome events.
Sign up free to answer and see why

Could you run a full modeling round — rapid-fire in 90s each, scenarios as axis→constraint→lever→verify, and the four derivations cold?

New to itGetting thereConfident

Takeaways

  • The round scores process over answer: open every question by naming the axis, the constraint, and the verification step.
  • Every answer has the shape mechanism → lever → verification; memorize the shape, not just the facts.
  • Scenarios reward sequencing and diagnosis; when a popular default exists (SMOTE, “calibrate everything,” ROC-AUC), say why it’s not your first move.
  • Be able to BOTH derive (bias-variance cancellation, logistic gradient/Hessian, MSE saturation) and explain it in words.
  • The universal red flag is the definition-only answer; convert each into a decision plus how you’d check it.

You’ve completed ML Fundamentals for Interviews — revisit any lesson’s checkpoints to keep the levers sharp before the real round.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.