Constraints first, metric second. The accuracy paradox, precision/recall/F1, ROC-AUC vs PR-AUC on imbalance and why the choice flips, thresholds via calibrated cost (not Youden), and the silent killer — data leakage, with the named case studies and the CV protocols that prevent it.
Metric selection is constraint-driven, not accuracy-driven
Two failure modes sink more offline-good/online-bad models than any modeling mistake: the wrong metric (accuracy on imbalanced data, ROC-AUC at the top of a rare-class list) and data leakage (the offline-AUC-0.95-then-online-0.62 surprise). The interview is explicit about both — DataInterview’s fraud prompt (“0.2% fraud, 500 reviews/day — which metric and threshold?”) and Stripe’s bug-squash round (“this model is too good to be true — find the leak”). The strong answer always starts from the business constraint (capacity, cost, latency), then picks the metric that lives at it. This lesson makes both reflexes automatic.
The confusion matrix & why accuracy lies
Every binary metric is a ratio of the four confusion-matrix cells — TP, FP, FN, TN — so fluency starts there. Precision = TP/(TP+FP) (of my alarms, how many are real), Recall = TP/(TP+FN) (of the real positives, how many I caught), FPR = FP/(FP+TN) (of the negatives, how many I falsely flagged). Notice precision and FPR have different denominators — precision divides by predicted positives, FPR by actual negatives — and that single difference is the entire ROC-vs-PR story below. Being able to write the four cells and derive each metric on demand is the floor an interviewer expects.
Start with why accuracy lies — the accuracy paradox. If the positive class is 1% of the data, “always predict negative” scores 99% accuracy and detects nothing. So accuracy near the base rate hides total model failure on any problem where the majority class is >90%. The rule: report accuracy only alongside precision, recall, and the positive base rate — and if you cannot beat the “always predict majority” baseline, drop it. Interview angle. “Your fraud model is 99.8% accurate” should make you ask “what’s the base rate?” before anything else; failing to is an instant junior tell.
Precision, recall, F1 — and the asymmetric-cost trap
The four-metric toolkit, each with its lying condition. Precision = of the alarms I raise, how many are real (lies if recall is tiny). Recall = of the real positives, how many I catch (lies if precision is tiny). F1 = harmonic mean of the two — punishes either being near zero, but hides asymmetric business cost: when a false negative costs $10,000 and a false positive costs $10, F1 is the wrong summary and you want expected cost. ROC-AUC = probability a random positive outranks a random negative (a ranking/discrimination measure, insensitive to calibration). PR-AUC = average precision across the PR curve (sensitive to base rate — only compare within a dataset).
The precision/recall tension is set by which error costs more, and the canonical drill (Damien Martin) is to ground it in the population. “1,000 patients, 1% prevalence, the test has recall 80% and precision 75% — if it flags 11 people, how many are actually sick?” Answer: precision 0.75 × 11 ≈ 8 true positives, ~3 false alarms, while recall 80% means it missed ~2 of the 10 sick. The grading signal is naming the cost asymmetry first — cancer screening favors recall (a missed case is fatal), inbox spam favors precision (a misfiled real email is worse than a stray spam) — then doing the size-of-denominator arithmetic instead of hand-waving “75% of the time it’s right.”
code
1THE CONFUSION MATRIX, AND EVERY METRIC FROM IT23 predicted + predicted -4 actual + TP FN <- recall = TP/(TP+FN)5 actual - FP TN <- FPR = FP/(FP+TN)6 ^precision = TP/(TP+FP)78 accuracy = (TP+TN)/all <- LIES under imbalance (~= base rate)9 F1 = 2PR/(P+R) <- harmonic mean; hides asymmetric FP/FN cost10 ROC-AUC : TPR vs FPR <- ranking; FPR denominator is huge when rare11 PR-AUC : precision vs recall<- precision denominator tracks the rare class
ROC-AUC vs PR-AUC: why the choice flips on imbalance
This is the most-asked metric question, and the mechanism is the whole answer (Saito & Rehmsmeier, PLOS ONE 2015, 6,300+ citations; Davis & Goadrich). ROC plots TPR vs FPR; when negatives vastly outnumber positives, even many false positives barely move FPR (its denominator is huge), so the curve can look great while precision at the top of the ranked list collapses. PR plots precision vs recall, and precision falls immediately as recall climbs past the base rate — surfacing exactly the failure ROC hides. So: ROC-AUC is fine for broad model comparison; PR-AUC is the honest deployment metric when the positive class is rare.
code
1WHICH CURVE TO TRUST (imbalanced positives)23 ROC plots TPR vs FPR. FPR = FP / (FP + TN).4 With TN huge (rare positives), FPR stays ~0 even with many FPs5 -> ROC-AUC looks optimistic; precision at top-K can be unusable.67 PR plots Precision vs Recall. Precision = TP / (TP + FP).8 Precision drops the instant recall exceeds the base rate9 -> PR-AUC tracks what you actually ship.1011 Heart-failure study: ROC-AUC 0.9451 but PR-AUC 0.911312 -> the model is better at the negatives than at the positives clinicians want.13 Rule: report BOTH; decide deployment on PR-AUC / Precision@K when positives are rare.
Interview angle. The capacity-constrained framing is where seniors separate: “0.2% fraud, ops reviews 500 alerts/day — which metric?” The answer is Precision@500 and Recall@500, set the threshold to yield exactly 500 alerts on recent validation data, and report it with confidence intervals — not ROC-AUC, which “can appear successful even when precision at the top of the list is unusable under extreme imbalance.” The metric must live at the actual decision point (the review budget), and the threshold is wired to the business KPI so it propagates as volume shifts. Naming the constraint before the metric is the rubric (Damien Martin, Dan Lee).
Thresholds: calibrated cost, not a ROC sweep
A common mistake is to scan the ROC curve for the point maximizing Youden’s J (sensitivity + specificity − 1) and ship that threshold. Manokhin’s practitioner argument: that is a “blunt instrument” when error costs are asymmetric or when downstream consumers want a real probability — and calibration “obviates the need for empirical ROC-based threshold hunting because it aligns scores with actual probabilities.” Once the model is calibrated (Lesson 3), the decision rule becomes principled: threshold at 0.5 for minimum-error with symmetric costs, or at the cost-optimal p* = C(FP) / (C(FP) + C(FN)) for asymmetric costs. So thresholding is a calibration + cost-matrix problem, not a curve-fitting exercise.
Class imbalance is handled in a strict order — and getting the order right is the rubric (Dan Lee, Travis Tang): (1) fix the metric away from accuracy (PR-AUC, recall at fixed precision, cost-weighted) — this changes only evaluation; (2) tune the threshold / class weights — never default to 0.5 under a skewed prior; (3) cost-sensitive loss (weighted cross-entropy, focal loss); (4) only then resample — SMOTE/ADASYN, with the caveat that they help when the minority has dense clusters and hurt when it is sparse noise (synthetic points amplify the noise). Class weights are usually preferred over resampling because they keep the data distribution intact. Interview angle. Leading with SMOTE is the classic weak answer; leading with “fix the metric, then the threshold” is the strong one.
code
1CLASS IMBALANCE: THE ORDER OF OPERATIONS23 1. METRIC -> PR-AUC / Recall@fixed-precision / cost-weighted (eval only)4 2. THRESHOLD -> tune to capacity or cost; never default 0.5 (no retrain)5 3. CLASS WTS -> weighted cross-entropy or focal loss (keeps data intact)6 4. RESAMPLE -> SMOTE/ADASYN LAST; helps dense minority,7 HURTS sparse/noisy minority (amplifies noise)89 Red flag in interviews: jumping straight to (4) "use SMOTE".10 Strong answer: diagnose the metric first, then the threshold.
One more metric family for when the output is a probabilistic forecast rather than a decision (demand, price, risk): log loss and the Brier score are proper scoring rules that penalize over- and under-confidence, so they reward calibrated probabilities directly. And when the output is a ranking with no fixed cutoff (search, recsys), the metric moves to position-aware measures — NDCG, MAP — that only need correct ordering, not calibrated p. Interview angle. “Which metric for a price/demand model?” → Brier/log loss, not accuracy; “for a search ranker?” → NDCG/MAP. Picking the metric to the output type (decision vs probability vs ranking) is half the battle.
A metric is a contract about what “good” means. Choose it from the constraint and the output type — capacity, cost, ranking, probability — before you ever look at a model’s score. Optimizing the wrong metric perfectly is the most expensive way to fail.
Data leakage: the silent inflator of offline metrics
Kaufman et al. (2012) define leakage as “the introduction of information about the target that should not legitimately be available to mine from” — a violation of the learn/predict separation, and the largest single source of offline-to-online collapse. The named forms to memorize: target leakage (a feature that encodes the label or a post-event fact — e.g. a fraud model using chargeback_received, which is only populated after a case is confirmed, so at serving time it’s null and the model has no signal despite offline AUC ≈ 1.0); preprocessing leakage (fitting a StandardScaler / target-encoder / TF-IDF on train+test before splitting); temporal leakage (random k-fold on time-ordered data lets the future leak backward); group leakage (the same entity spans train and test); and duplicate-row leakage.
code
1LEAKAGE TAXONOMY (memorize the named forms + detection)23 form canonical example detection4 --------------- ------------------------------------- --------------------------5 target fraud model uses chargeback_received "would this column exist6 (set only AFTER the label) at prediction time?"7 preprocessing StandardScaler fit on train+test fit() only on train rows8 temporal random k-fold on daily sales forward-chaining CV9 group/subject same patient in train AND test GroupKFold / leave-one-subject10 duplicate-row near-dup rows across the split hash-based dedup across folds1112 Measured inflation (connectome ML, Rosenblatt 2024 Nature Comms):13 attention prediction r = 0.01 (true) -> r = 0.48 with feature leakage.14 Smaller samples are MORE susceptible.
The case studies give you numbers to cite. Feature-selection leakage (Rosenblatt et al., Nature Communications 2024, ABCD connectomes): selecting features on combined train+test inflated attention-problem prediction from chance (r=0.01) to r=0.48 — a useless model reported as strong — and “smaller samples are more susceptible.” Group leakage (Lee et al. 2023, EEG PTSD): trial-wise CV that splits one patient’s augmented trials across train and test “yielded significantly inflated performance” vs subject-wise CV, because the model recognized the individual, not the pathology. The fixes are protocol, not modeling: feature selection inside the CV loop, forward-chaining for time series, GroupKFold/leave-one-group-out for recurring entities, and every preprocessing fit() on training rows only.
Interview angle. Stripe’s bug-squash: “the model is too good to be true — find the leak.” The methodical answer (which is what they grade): trace each feature to its source timestamp and ask “would this be populated at prediction time?”; flag any aggregation whose window includes the prediction timestamp; audit scaling/encoding statistics for train+test contamination; check join keys for sneak-into-train rows; then isolate by removing the single highest-importance feature and watching the score collapse — an abnormally important feature for an outcome that shouldn’t be that predictable is leakage until proven otherwise. The weak answer (“add regularization and see if the score drops”) treats a data-integrity bug as a modeling knob.
Interview prep
Metrics-and-leakage questions reward starting from the constraint and the data-generating process. Define the business limit, pick the metric that lives at it, set the threshold from calibrated cost, and prove the split prevents leakage. Have these ready.
01“When does accuracy lie?” → imbalanced data: predicting the majority hits the base rate (99% on 1% positives) and detects nothing; report it only with precision/recall/base-rate.
02“ROC-AUC vs PR-AUC?” → ROC for broad comparison; PR-AUC (and Precision@K) for rare positives, because FPR’s huge denominator hides precision collapse at the top.
03“0.2% fraud, 500 reviews/day — metric?” → Precision@500 / Recall@500; set the threshold to yield 500 alerts on recent validation, not a ROC sweep.
04“How do you set a threshold?” → calibrate, then cut at 0.5 (symmetric) or p* = C(FP)/(C(FP)+C(FN)); don’t ship Youden’s J under asymmetric cost.
05“Class imbalance — first move?” → fix the metric, then threshold/class-weights, then cost-sensitive loss, and only then resample (SMOTE can amplify noise).
06“Most common leakage in time series?” → random k-fold leaks the future; use forward-chaining and only features observable at prediction time t.
07“Same user in train and test?” → group leakage; use GroupKFold / leave-one-group-out so no entity ID crosses the split.
08“Model looks too good — debug?” → trace feature timestamps, audit preprocessing fit-on-train-only, remove the top feature and watch the score collapse.
Follow-ups: “F1 vs precision@K?” (F1 hides asymmetric cost and the PR frontier; precision@K operates at the real cutoff), “why is AUC high but online engagement dropped?” (verify evaluation parity — same dataset version, label definition, joins — then backtest old vs new through each pipeline to isolate drift/label-delay), and the medical-CV curveball (“random k-fold gave 0.95 — trust it?” → no, check for subject leakage with subject-wise CV). The deeper signal is hypothesis-driven debugging (Stripe’s rubric) and grounding the metric in the population (Damien Martin’s 1000-patient drill: of 11 flagged at precision 0.75, ~8 truly positive).
A churn model is 97% accurate, but only 3% of customers churn and the model catches almost no churners. In an interview, what’s the best next step?
ASwitch to imbalance-aware metrics (PR-AUC, recall at fixed precision, cost-based), then tune the threshold and consider class weights — accuracy here is meaninglessBReport the 97% accuracy as a strong resultCImmediately apply SMOTE to balance the classes and retrain
Two fraud models have nearly identical ROC-AUC (~0.97), but ops can only review the top 300 alerts/day. How do you choose between them?
APick whichever has the higher ROC-AUC — it’s the standard comparisonBCompare Precision@300 (and Recall@300) on recent validation data — the metric must live at the actual review capacity, where ROC-AUC is blindCChoose the model with the lower log loss overall
An EEG model classifying patients hits 0.96 accuracy with random 5-fold CV. Each patient contributes ~40 augmented trials. What should you suspect and do?
ATrust it — 0.96 with cross-validation is robustBLower the learning rate and retrain to confirm the accuracy is stableCSuspect group (subject) leakage; re-evaluate with subject-wise CV (no patient in both train and test) — expect the inflated accuracy to drop
A teammate standardizes the full dataset with StandardScaler, then splits into train/test and reports a suspiciously strong score. What’s the bug and the fix?
ANo bug — scaling before splitting is standard practiceBPreprocessing leakage: the scaler learned mean/variance from train+test; fit the scaler inside the CV loop / Pipeline on training rows only, then transform testCThe test set is too small — enlarge it to stabilize the score
Your binary classifier is well-calibrated. A false negative costs $1,000; a false positive costs $50. Where should the decision threshold go, and why not Youden’s J?
AAt 0.5, because that’s the natural cutoff for a calibrated modelBAt the ROC point maximizing Youden’s J (sensitivity + specificity − 1)CAt p* = C(FP)/(C(FP)+C(FN)) = 50/1050 ≈ 0.048 — a calibrated probability lets you cut at the cost-optimal point
Could you pick the metric from the constraint, defend PR-AUC vs ROC-AUC on imbalance, set a cost-based threshold, and find/prevent leakage with the right CV?
New to itGetting thereConfident
Takeaways
Constraint first, metric second: accuracy lies on imbalance (it’s the base rate); report it only with precision/recall/base-rate.
ROC-AUC for broad comparison; PR-AUC / Precision@K when positives are rare — FPR’s huge denominator hides precision collapse.
Threshold = calibration + cost matrix: cut at p* = C(FP)/(C(FP)+C(FN)), not a Youden sweep.
Imbalance order: metric → threshold/class-weights → cost-sensitive loss → resample (SMOTE last; it can amplify noise).
Leakage is the silent killer: time-aware + group-aware splits, fit preprocessing on train only, feature-select inside CV; remove the top feature to expose it.
Next: the timed mock — a rapid-fire gauntlet across the fundamentals, under interview pressure.