Lesson 4 of 4 · 35 min

Present a testable ML plan in ten minutes

Write a concise decision record with a baseline, test, and stop rule.

A strong ML interview answer makes the next decision easier. Start by restating the target action and the population. Then name the simplest baseline that can establish whether the data and evaluation work. Add complexity only when a specific failure analysis suggests it will help.
A baseline is an instrument for learning. A constant prediction checks the metric and prevalence. A simple heuristic checks whether readily available information already solves much of the task. A small linear or tree model can expose label and feature problems before a large model consumes time. A sophisticated model without these comparisons may hide an avoidable data defect.
State the evaluation design before the result. Identify the split unit, time cutoff, label horizon, primary metric, operational constraints, and test data that remains untouched. Name a plausible failure slice and a guardrail. Give a resource budget and a stopping rule. This turns "we will try several models" into a falsifiable plan.
Discuss alternatives with evidence. If the baseline misses nonlinear interactions, test a model that can represent them. If it fails on recent users because labels are immature, model complexity is not the first remedy. If a feature is unavailable online, remove or replace it before comparing architectures. Good judgment includes rejecting an experiment whose result would not answer the question.

Worked example

A fictional platform wants to prioritize support tickets that will need specialist escalation within 24 hours. The plan begins with a rule based on issue category and account status, then compares a calibrated linear classifier. Training uses historical features available at ticket creation. Labels mature after the 24-hour window plus ingestion allowance.
Validation uses later weeks, with a separate new-account slice. The operating metric is recall at a maximum of 200 reviews per day, with precision and latency reported. The first experiment has a fixed two-day analysis budget and stops if feature availability cannot be reproduced. The release decision requires a measurable improvement at the same review capacity and no known input-contract failure.

Exercise and solution

Design a one-page plan for predicting whether a user will complete onboarding within seven days. You have event logs, device type, signup time, and eventual completion time.
A model solution defines prediction at signup, excludes later onboarding events from initial features, waits for labels to mature, uses a temporal holdout, and starts with a prevalence baseline plus simple model. It reports a metric tied to the intended intervention and notes that predicting completion does not prove an intervention causes it. Award one point each for cutoff, label horizon, baseline, valid split, and decision-specific evaluation. Extra credit requires a proposed experiment for the intervention itself, kept separate from predictive validation.

Lab artifact: a one-page decision record

The record below is an original planning example. It makes the proposed model answer a decision rather than treating modeling as the goal.
code
1Action: offer optional human setup help at signup2Prediction target: onboarding completion within seven days3Population: new organization administrators in the supported signup flow4Features: values available at signup, with availability audit5Primary predictive comparison: review yield at 100 offers per day6Baselines: prior prevalence; current help-request rule; small linear model7Development split: earlier cohorts / later mature cohorts8Final evaluation: held-out later cohort, fixed once before opening labels9Guardrail: no unsupported input route; scoring p95 below 80 ms10Budget: two days for data audit; three candidate training runs11Stop: invalid label collection or unreproducible online features12Next causal question: does offering help improve completion?
The capacity metric describes prediction-based selection. It does not answer the intervention question. A model can identify users unlikely to finish who also cannot benefit from the offered help. Conversely, a moderately risky group might respond strongly. The causal follow-up needs a defined treatment, comparison, assignment unit, outcome window, and decision criteria. Keep those questions connected but separate.

A second failure case: the baseline reveals an identifier shortcut

Suppose a linear model and a large model both achieve ninety-nine percent validation accuracy. The simplest prevalence baseline is seventy percent. Before declaring success, inspect suspicious features. An account identifier includes a prefix assigned only after successful onboarding, and the training extraction mistakenly joins the latest account record rather than its signup state.
Removing that feature reduces both models to similar performance. The initial model comparison did not show that both learned the task well; both exploited the same unavailable information. A more complex architecture would not fix the data join. The decisive artifact is a timeline showing when the identifier prefix becomes available, followed by a replay using the signup-time state.
CheckIf it failsNext action
Feature existed at signupFuture information enters predictionRepair extraction
Outcome window completeUnknowns look like negativesRebuild eligible cohort
Training fixture behaves correctlyLearning mechanics uncertainDebug local update
Baseline meets action requirementComplexity may add no needed valueStop or test a named gap
Candidate improves fixed policyPotential predictive gainValidate live compatibility

Exercise: allocate a limited week

You have five working days. The data audit requires two days, baseline evaluation one day, and each large-model run one day. There is a known unresolved question about whether outcome logging covers mobile users. Choose a plan and state when it stops.
Reserve the first two days for the data and label audit, including mobile coverage. If the labels cannot be made valid within the budget, stop the model comparison and deliver the audited gap with affected scope. If the audit passes, spend one day on the baseline and error analysis. Use remaining time for a targeted candidate only if a specific gap justifies it, retaining time to inspect the result and write the decision. Starting two large runs immediately would consume time without resolving whether their scores mean anything. Award one point each for dependency order, stop condition, baseline, targeted hypothesis, and review time.

Misconceptions to correct

“A baseline is only a number to beat” misses its diagnostic role. A baseline can reveal a prevalence mismatch, leakage, or a pipeline defect. “A good risk predictor tells us whom an intervention will help most” confuses predicted outcome under historical conditions with treatment benefit. The latter depends on a comparison of possible actions, not only a risk score.
In the interview, state what result would make you stop. If a simple policy meets the requirement with lower maintenance burden and no identified failure gap, stopping is a reasoned outcome. If the candidate misses the latency limit, a higher offline metric does not erase that constraint. Your final record should identify the next owner, the exact unresolved question, and the evidence that would reopen the decision. This makes the plan usable by another engineer after the interview ends.

Interview probe

Original practice: You have one week to improve an ML product. What do you do first? A strong answer chooses a decision, validates data, measures a baseline, and runs a bounded test informed by errors. Follow up with a baseline that already meets the requirement. A weak answer lists architectures without a stopping rule.

Sources

docsGoogle: rules of machine learningdevelopers.google.comdocsAmazon: Applied Scientist interview preparationamazon.jobsdocsMicrosoft Research: checks after an online experimentmicrosoft.com

Checkpoint

Which first plan makes the modeling decision testable?

ATune architectures on the final holdout until capacity fits.BDefine baseline, valid split, fixed operating metric and stop rule.CChoose the success metric after error analysis.DRun the largest candidate before checking labels.
Sign up free to answer and see why

Checkpoint

An account prefix is assigned after onboarding. Can a signup-time predictor use the latest prefix?

AYes, because it is stored in the same account record.BYes, if a simpler model also uses it.CNo; it was unavailable at the prediction cutoff.DOnly if the final test accuracy is high.
Sign up free to answer and see why

Checkpoint

A simple baseline already meets the action requirement with no identified error gap. What is a defensible next decision?

AStop or propose a specific additional gap before adding complexity.BTrain a large model solely because time remains.CRedefine the requirement to ensure the baseline fails.DReuse test labels until a candidate wins.
Sign up free to answer and see why

Checkpoint

Mobile outcome logging is unresolved. Which sequence uses a limited week well?

ARun candidates first and audit only the best result.BTreat missing mobile outcomes as negatives for speed.CRemove mobile users after seeing which improves the score.DAudit the outcome scope, then baseline, then a justified candidate.
Sign up free to answer and see why

Checkpoint

A model identifies users unlikely to complete onboarding. What is still needed to claim offered help improves completion?

AHigher predictive AUC alone.BA valid intervention comparison with outcome follow-up.CA feature importance ranking.DA lower training loss on those users.
Sign up free to answer and see why

Can you give another engineer a baseline, valid target, budget, and stop rule, then separate prediction from intervention? Rate confidence from 1 to 5 and name the first unresolved dependency.

Not yetGetting thereConfident

Wrap-up

  • An interview plan should state what result would change your decision. Make the baseline and evaluation executable on the available data.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.