The decision is a function of data shape, volume, latency, interpretability, and cost — not a leaderboard. Logistic regression vs gradient-boosted trees vs neural nets vs LLMs, with named case studies where the “smaller” model won, and the real cost/latency numbers.
A decision framework, not a leaderboard
The senior mistake is to reach for the highest-capacity model and tune it. The senior skill is to choose by constraint: data shape (tabular vs unstructured), data volume, latency, interpretability burden, inference cost, and maintenance posture. Chip Huyen’s framing (Designing ML Systems) is the through-line: pick the cheapest model that clears a clear eval gate, then pay for complexity only when the metrics demand it. The interview tests whether you start from the constraint — Stripe’s modeling rubric is literally “articulate trade-offs like accuracy vs processing speed vs regulatory compliance,” not “name the best model.”
The two highest-leverage axes are data shape and latency; everything else is a downstream constraint. Tabular/structured → tree ensembles (XGBoost/LightGBM/CatBoost) and a logistic baseline. Unstructured (text, images, audio) → neural nets, because they exploit a structural prior (translation invariance, autoregressive dependence) that a generic tabular learner cannot. LLMs sit on top when the task is open-ended language/reasoning or labels are scarce — they are not a drop-in replacement for supervised tabular models. Memorize that decomposition; it is the spine of every answer.
code
1MODEL FAMILIES AT A GLANCE (senior decision table)23 axis Logistic Reg GBT (XGB/LGBM) Neural Net LLM (GPT-4-class)4 -------------- -------------- ---------------- --------------- ------------------5 wins at <10k, linear 10k-10M tabular >100k + unstruct scarce labels, text6 interpretability highest medium (SHAP) low lowest7 train cost seconds minutes-hours hours-days $0 prompt / $1k-1M FT8 latency/sample <1 ms 1-10 ms 10-50 ms 300-3000 ms (out-bound)9 inference cost ~free ~free GPU $ $0.62-24 / 10k tokens10 fails by underfit OOD categories drift, forgetting hallucination, fragility1112 Two axes dominate: DATA SHAPE and LATENCY. The rest are downstream constraints.
code
1THE DECISION TREE (walk it out loud in the interview)23 Is the data tabular/structured?4 YES -> small & low-noise/autocorrelated? -> seasonal-naive / ARIMA / logistic baseline5 -> otherwise -> GRADIENT-BOOSTED TREES (default)6 -> need per-decision explanation? -> logistic regression (ship if AUC gap small)7 NO (text / image / audio)?8 -> have labels + latency budget? -> neural net (or fine-tuned small model)9 -> labels scarce / open-ended language -> LLM (layer it; don't serve on a hot path)1011 Then: is it latency-critical (high QPS, tight p99)? -> stay on trees/linear, not an LLM.12 Always: train a CALIBRATED baseline first; climb capacity only on measured lift.
GBTs are “the linear-regression-of-the-2020s” for tabular — the baseline everything else must beat. Three structural reasons (and they make good interview ammo): (1) scale-invariance — trees split on thresholds, so a column in dollars and one in pennies coexist with no scaling; (2) native missing-value handling — XGBoost/LightGBM learn a default direction per split, so sparse/absent categories rarely need imputation; (3) interactions for free — a depth-6 tree encodes three-way interactions you never wrote down. The r/datascience production consensus: XGBoost has “sufficient flexibility to model complex patterns without the same overfitting risk” as deep nets, with fewer fragile hyperparameters. The follow-on survey “Tabular Data: Is Deep Learning All You Need?” (arXiv 2402.03970) confirms the pattern holds for newer architectures: TabM “trails the top methods by a narrow margin, ranking third,” with CatBoost/XGBoost clustered at the top.
The failure modes a senior plans for: out-of-distribution categorical levels (a brand-new merchant category routes to the default branch, often poorly — mitigate with a new-level detector that falls back to a rule or a logistic sub-model); feature decay (serving distribution drifts from training — retrain weekly on a sliding window, monitor per-feature KS drift); and interaction overfit (branching on noisy three-way interactions — cap max_depth at 6–8, early-stop on a holdout, sanity-check SHAP). Interview angle. “Why does XGBoost beat a neural net on our tabular data?” → scale-invariance + missing-value handling + cheap interactions + fewer fragile knobs, and you only need >1M rows or genuine unstructured signal to justify the NN R&D budget.
When the “smaller” model wins — named case studies
The deep research surfaces three cases where capacity lost, and they are exactly what “bias-variance + No-Free-Lunch” predicts. (1) Seasonal Random Walk beats everything on EPS forecasting (Kurylek 2024): a naive seasonal no-change benchmark beat MLP, CNN, LSTM and XGBoost on quarterly Polish EPS, statistically significant by Wilcoxon — because the data-generating process was too simple for any nonlinear learner, so their higher VC dimension only bought variance. (2) ARIMA beats XGBoost on COVID-19 forecasting (Rahman 2022) — same lesson on small, autocorrelated series. (3) Logistic regression matches XGBoost on fraud (Rozaq 2025): LR hit the highest recall (70.45%) and, with feature engineering, AUC rose 0.799 → 0.909 — and it shipped because regulated fraud-ops need to explain a coefficient, not a tree ensemble.
code
1WHEN CAPACITY LOST (the cases interviewers love)23 problem (small / structured) winner loser(s)4 ------------------------------- --------------- ---------------------------5 Quarterly EPS forecast Seasonal RW MLP, CNN, LSTM, XGBoost6 COVID-19 case forecast ARIMA XGBoost7 Regulated fraud (interpretable) Logistic Reg XGBoost (tied AUC, LR shipped)89 Lesson: NOT "simple always wins" -- winning is dataset-dependent.10 Small + autocorrelated + low-noise -> no-capacity baseline.11 Always benchmark a CALIBRATED baseline before training a tree or net.
The takeaway is not “simple always wins” — it is “winning is dataset-dependent, so benchmark a calibrated baseline first.” The benchmark order that signals seniority: seasonal-naive / ARIMA / logistic → GBT → neural net → foundation model, climbing only when the residual analysis shows signal left on the table and the lift clears the maintenance cost (a common bar: ship the simpler model unless the AUC gap exceeds ~0.03–0.05 and the regulator accepts post-hoc explanation). Interview angle. “When would you not use deep learning?” → small/low-noise/autocorrelated data, regulated interpretability requirements, latency-critical hot paths, or when a calibrated GBT already clears the eval gate.
Interview angle. “But you can explain a GBT with SHAP — doesn’t that close the interpretability gap?” This is a sharp follow-up worth a careful answer: SHAP gives post-hoc, local attributions that are excellent for debugging and for satisfying many regulators, but they are an approximation of a black box, not the model’s actual decision rule — and a regulator who requires a monotonic, auditable relationship (e.g. “income up ⇒ risk never increases”) may not accept them. So the honest framing is: SHAP narrows the gap for explanation-to-engineers and many compliance regimes, but a strictly interpretable model (logistic regression, or a monotonicity-constrained GBT) is still the safer ship when the relationship itself must be guaranteed.
LLMs: a different tool, priced by output tokens
LLMs are general-purpose function approximators over language/code — not a substitute for supervised tabular learning. They win on open-ended generation, summarization, long-context reasoning, messy-text entity extraction, and zero-/few-shot tasks where labels are scarce. Chip Huyen’s heuristics are interview-grade: “prompt engineering is a cheap and fast way to get something running,” and as examples accumulate “fine-tuning generally yields better performance,” with the rough rule “a prompt is worth ~100 examples.” The cost reality is dominated by inference, not training: GPT-4 at 10k-in/200-out ≈ $0.62/request vs GPT-3.5 at $0.004 — and “input length has little effect (parallel), but output length significantly increases latency because tokens are generated sequentially.” So text products budget for worst-case output length.
Where LLMs should not be first choice: high-QPS real-time scoring on structured features, low-latency ranking, any task needing a calibrated probability of a rare class, or a regulated workflow demanding per-decision explanation. The senior pattern is to layer LLMs on a supervised core — extracting features from free text, summarizing decisions for users, generating counterfactual explanations — rather than replacing a calibrated tabular model. Mitigations when you do use them: fine-tune a smaller open model to bake in instructions and shorten prompts, quantize (INT8 roughly halves KV-cache memory), cache repeated prefixes, and cap output length at the API layer. And note the different failure surface — prompt fragility, version churn when the provider updates the model, and hallucination (no calibrated score), which is why LLM apps need prompt unit tests and a held-out offline eval set.
Latency & cost at scale — the numbers that decide it
The constraint that most often kills the “bigger model” is the hot path, and seniors carry rough numbers. A logistic regression or GBT scores in under 1–10 ms on CPU at essentially zero marginal cost; a small neural net is 10–50 ms on CPU (single-digit ms on GPU); a frontier LLM is 300–3000 ms and dominated by sequential output tokens. The Databricks inference benchmarks make the LLM economics concrete: tensor-parallelism and hardware choice swing latency 12–52% (Llama2-70B on 4× H100 is ~36% lower latency at batch 1 than A100), and INT8 quantization roughly halves KV-cache memory (Llama2-70B at batch 32: 10 GiB → 5 GiB). None of that brings an LLM near a 20 ms p99 budget — which is why high-QPS scoring stays on trees.
There is also a perceived-quality argument that lands in product-minded rounds. Notion cut chat latency ~4× (from ~2s to ~350ms) for 100M+ users by serving a smaller fine-tuned model on dedicated infra rather than a frontier model on the hot path — their framing, “latency is perceived as quality,” is worth quoting. The lesson for model selection: the model that shrinks what hits the expensive component often wins on speed, cost, and user-perceived quality simultaneously, even at a tiny offline-accuracy cost.
The leaderboard rewards the last decimal of accuracy; production rewards the model you can serve at p99 under budget, retrain in a notebook, and explain to a regulator. Senior model selection optimizes the second list.
Cross-validation that matches deployment
Model selection is only as honest as the CV protocol underneath it, which is the bridge to Lesson 5. Nested CV (outer loop for evaluation, inner loop for tuning) prevents the optimism of tuning and evaluating on the same split. TimeSeriesSplit / forward-chaining is mandatory on time-ordered data — random k-fold leaks the future. GroupKFold / leave-one-group-out is mandatory when an entity (user, patient, device) recurs across rows. The rule: write down the data-generating process before you write the CV loop, and match the protocol to the deployment distribution — not to what a Kaggle leaderboard rewards. A model that wins under the wrong CV wins nothing in production.
Interview angle. “How do you tune hyperparameters without overfitting the validation set?” The honest answer names the optimism problem: every tuning trial peeks at the validation fold, so picking the best of 200 configs and reporting its validation score is upward-biased. The fix is a clean separation — tune on inner folds, report on an outer fold (nested CV) or a final untouched test set — and a search strategy matched to budget: grid search for a few knobs, random search or Bayesian optimization (and early-stopping pruners like Hyperband) when the space is large. Mentioning that random search beats grid search in high dimensions (Bergstra & Bengio) is a small senior flourish.
Interview prep
Model-selection questions are “show-your-prioritization” prompts. Start from the constraint (data shape, volume, latency, interpretability, cost), name a baseline, and justify any climb in capacity with a measured lift. Have these ready.
01“How do you choose a model?” → by data shape + latency first, then volume/interpretability/cost; cheapest family that clears the eval gate (Huyen).
03“When does a simpler model win?” → small, low-noise, autocorrelated data (Seasonal RW > XGBoost on EPS; ARIMA > XGBoost on COVID); regulated interpretability (LR ships over XGBoost).
04“When would you NOT use deep learning?” → small/clean tabular, latency-critical hot paths, calibration-critical decisions, regulated explanation needs.
05“When is an LLM the right tool?” → open-ended text/reasoning, scarce labels; layer it on a supervised core, don’t replace tabular models.
06“What dominates LLM cost/latency?” → output tokens (sequential decode); a prompt ≈ 100 examples, so fine-tune as data grows.
07“How big a lift justifies the complex model?” → typically >0.03–0.05 AUC AND worth the maintenance/latency tax; otherwise ship the simpler one.
08“What CV proves the choice?” → nested CV for tuning honesty; forward-chaining for time series; GroupKFold for recurring entities — match deployment.
Follow-ups push on tension: “XGBoost wins on tabular — always?” (no: the EPS/COVID counterexamples; benchmark a calibrated baseline), “LLM vs supervised — same axis?” (no: LLMs bootstrap cheaply then fine-tune; top Kaggle tabular solutions are tree boosters and Transformer/TabNet hybrids, not standalone LLMs), and the regulated-domain curveball (“the GBT is 0.02 AUC better but you can’t explain a decline to a customer” → ship the logistic model). The grading signal (Stripe, Exponent) is sequencing trade-offs and tying the choice to the problem framing, never naming a single “best” model.
You’re forecasting a quarterly business metric with ~120 historical points and strong seasonality. A teammate wants to start with an LSTM. What do you argue for?
AStart with a seasonal-naive / ARIMA baseline; only move to capacity if its residuals show real signal — small autocorrelated data often beats deep nets with a no-change benchmarkBAgree — LSTMs are state-of-the-art for sequences, so it’s the safe defaultCUse an LLM to read the time series as text and predict the next value
A regulated credit-risk model: a tuned XGBoost scores 0.02 AUC above your logistic regression, but compliance must explain every decline to the applicant. Which do you ship?
AXGBoost — always ship the highest AUCBNeither — collect more data until a neural net clearly winsCThe logistic regression — its coefficients give per-decision explanations the regulator accepts, and the AUC gap is below the bar that would justify the interpretability loss
You must classify 50,000 free-text support tickets/second into 8 categories with a p99 budget of 20 ms and no labeled data yet. What’s the most defensible first architecture?
ACall a frontier LLM per ticket with a zero-shot prompt and ship itBUse the LLM offline to bootstrap labels, then train a small, fast text classifier (or fine-tune a small model) that meets the 20 ms p99 — layer the LLM, don’t serve it on the hot pathCTrain a 70B model from scratch on the tickets for best accuracy
Your tabular churn model is a logistic regression at 0.82 AUC; a colleague’s deep net hits 0.83 but needs a GPU to serve and drifts monthly. Mixed categorical features have many missing values. Best next move?
AShip the deep net — 0.83 > 0.82BTry a gradient-boosted tree, which natively handles the missing values and interactions and usually tops both on tabular at near-zero serving cost — then compare honestlyCImpute all missing values with column means and re-run the logistic regression
An LLM text feature looks great offline but your cost projection is 5× budget at launch volume. Which lever set most directly fixes cost without abandoning the LLM?
AIncrease the temperature so the model writes more confidently and stops soonerBMove to a reasoning model so it thinks more per requestCFine-tune a smaller model to shorten prompts, cap output length at the API layer, cache repeated prefixes, and quantize — output tokens dominate cost
Could you choose a model from data shape + latency + interpretability + cost, defend a calibrated baseline, cite a case where capacity lost, and say when to layer an LLM?
New to itGetting thereConfident
Takeaways
Choose by data shape + latency first, then volume/interpretability/cost — cheapest family that clears the eval gate, not the leaderboard winner.
GBTs own tabular (scale-invariance, missing-value handling, free interactions); beat NNs unless >1M rows or genuine unstructured signal.
Capacity can lose: Seasonal RW > four learners on EPS; ARIMA > XGBoost on COVID; logistic regression ships over XGBoost in regulated fraud.
LLMs are for open-ended text / scarce labels — layer them on a supervised core; output tokens dominate cost and latency.
Match CV to deployment (nested / forward-chaining / GroupKFold); a win under the wrong CV is no win.
Next: evaluation metrics done right — precision/recall, ROC-AUC vs PR-AUC, thresholds, imbalance, and data leakage.