Write a concise decision record with a baseline, test, and stop rule.
A strong ML interview answer makes the next decision easier. Start by restating the target action and the population. Then name the simplest baseline that can establish whether the data and evaluation work. Add complexity only when a specific failure analysis suggests it will help.
A baseline is an instrument for learning. A constant prediction checks the metric and prevalence. A simple heuristic checks whether readily available information already solves much of the task. A small linear or tree model can expose label and feature problems before a large model consumes time. A sophisticated model without these comparisons may hide an avoidable data defect.
State the evaluation design before the result. Identify the split unit, time cutoff, label horizon, primary metric, operational constraints, and test data that remains untouched. Name a plausible failure slice and a guardrail. Give a resource budget and a stopping rule. This turns "we will try several models" into a falsifiable plan.
Discuss alternatives with evidence. If the baseline misses nonlinear interactions, test a model that can represent them. If it fails on recent users because labels are immature, model complexity is not the first remedy. If a feature is unavailable online, remove or replace it before comparing architectures. Good judgment includes rejecting an experiment whose result would not answer the question.
Worked example
A fictional platform wants to prioritize support tickets that will need specialist escalation within 24 hours. The plan begins with a rule based on issue category and account status, then compares a calibrated linear classifier. Training uses historical features available at ticket creation. Labels mature after the 24-hour window plus ingestion allowance.
Validation uses later weeks, with a separate new-account slice. The operating metric is recall at a maximum of 200 reviews per day, with precision and latency reported. The first experiment has a fixed two-day analysis budget and stops if feature availability cannot be reproduced. The release decision requires a measurable improvement at the same review capacity and no known input-contract failure.
Exercise and solution
Design a one-page plan for predicting whether a user will complete onboarding within seven days. You have event logs, device type, signup time, and eventual completion time.
A model solution defines prediction at signup, excludes later onboarding events from initial features, waits for labels to mature, uses a temporal holdout, and starts with a prevalence baseline plus simple model. It reports a metric tied to the intended intervention and notes that predicting completion does not prove an intervention causes it. Award one point each for cutoff, label horizon, baseline, valid split, and decision-specific evaluation. Extra credit requires a proposed experiment for the intervention itself, kept separate from predictive validation.
Lab artifact: a one-page decision record
The record below is an original planning example. It makes the proposed model answer a decision rather than treating modeling as the goal.
code
1Action: offer optional human setup help at signup2Prediction target: onboarding completion within seven days3Population: new organization administrators in the supported signup flow4Features: values available at signup, with availability audit5Primary predictive comparison: review yield at 100 offers per day6Baselines: prior prevalence; current help-request rule; small linear model7Development split: earlier cohorts / later mature cohorts8Final evaluation: held-out later cohort, fixed once before opening labels9Guardrail: no unsupported input route; scoring p95 below 80 ms10Budget: two days for data audit; three candidate training runs11Stop: invalid label collection or unreproducible online features12Next causal question: does offering help improve completion?
The capacity metric describes prediction-based selection. It does not answer the intervention question. A model can identify users unlikely to finish who also cannot benefit from the offered help. Conversely, a moderately risky group might respond strongly. The causal follow-up needs a defined treatment, comparison, assignment unit, outcome window, and decision criteria. Keep those questions connected but separate.
A second failure case: the baseline reveals an identifier shortcut
Suppose a linear model and a large model both achieve ninety-nine percent validation accuracy. The simplest prevalence baseline is seventy percent. Before declaring success, inspect suspicious features. An account identifier includes a prefix assigned only after successful onboarding, and the training extraction mistakenly joins the latest account record rather than its signup state.
Removing that feature reduces both models to similar performance. The initial model comparison did not show that both learned the task well; both exploited the same unavailable information. A more complex architecture would not fix the data join. The decisive artifact is a timeline showing when the identifier prefix becomes available, followed by a replay using the signup-time state.
Check
If it fails
Next action
Feature existed at signup
Future information enters prediction
Repair extraction
Outcome window complete
Unknowns look like negatives
Rebuild eligible cohort
Training fixture behaves correctly
Learning mechanics uncertain
Debug local update
Baseline meets action requirement
Complexity may add no needed value
Stop or test a named gap
Candidate improves fixed policy
Potential predictive gain
Validate live compatibility
Exercise: allocate a limited week
You have five working days. The data audit requires two days, baseline evaluation one day, and each large-model run one day. There is a known unresolved question about whether outcome logging covers mobile users. Choose a plan and state when it stops.
Reserve the first two days for the data and label audit, including mobile coverage. If the labels cannot be made valid within the budget, stop the model comparison and deliver the audited gap with affected scope. If the audit passes, spend one day on the baseline and error analysis. Use remaining time for a targeted candidate only if a specific gap justifies it, retaining time to inspect the result and write the decision. Starting two large runs immediately would consume time without resolving whether their scores mean anything. Award one point each for dependency order, stop condition, baseline, targeted hypothesis, and review time.
Misconceptions to correct
“A baseline is only a number to beat” misses its diagnostic role. A baseline can reveal a prevalence mismatch, leakage, or a pipeline defect. “A good risk predictor tells us whom an intervention will help most” confuses predicted outcome under historical conditions with treatment benefit. The latter depends on a comparison of possible actions, not only a risk score.
In the interview, state what result would make you stop. If a simple policy meets the requirement with lower maintenance burden and no identified failure gap, stopping is a reasoned outcome. If the candidate misses the latency limit, a higher offline metric does not erase that constraint. Your final record should identify the next owner, the exact unresolved question, and the evidence that would reopen the decision. This makes the plan usable by another engineer after the interview ends.
Interview probe
Original practice: You have one week to improve an ML product. What do you do first? A strong answer chooses a decision, validates data, measures a baseline, and runs a bounded test informed by errors. Follow up with a baseline that already meets the requirement. A weak answer lists architectures without a stopping rule.
Which first plan makes the modeling decision testable?
ATune architectures on the final holdout until capacity fits.BDefine baseline, valid split, fixed operating metric and stop rule.CChoose the success metric after error analysis.DRun the largest candidate before checking labels.
An account prefix is assigned after onboarding. Can a signup-time predictor use the latest prefix?
AYes, because it is stored in the same account record.BYes, if a simpler model also uses it.CNo; it was unavailable at the prediction cutoff.DOnly if the final test accuracy is high.
A simple baseline already meets the action requirement with no identified error gap. What is a defensible next decision?
AStop or propose a specific additional gap before adding complexity.BTrain a large model solely because time remains.CRedefine the requirement to ensure the baseline fails.DReuse test labels until a candidate wins.
Mobile outcome logging is unresolved. Which sequence uses a limited week well?
ARun candidates first and audit only the best result.BTreat missing mobile outcomes as negatives for speed.CRemove mobile users after seeing which improves the score.DAudit the outcome scope, then baseline, then a justified candidate.
A model identifies users unlikely to complete onboarding. What is still needed to claim offered help improves completion?
AHigher predictive AUC alone.BA valid intervention comparison with outcome follow-up.CA feature importance ranking.DA lower training loss on those users.
Can you give another engineer a baseline, valid target, budget, and stop rule, then separate prediction from intervention? Rate confidence from 1 to 5 and name the first unresolved dependency.
Not yetGetting thereConfident
Wrap-up
An interview plan should state what result would change your decision. Make the baseline and evaluation executable on the available data.