Specify a candidate rollout with a fallback and measurable stop conditions.
Offline validation evaluates a model on recorded examples. A release places the model inside a live decision system. The system adds request parsing, feature retrieval, latency, missing data, and user feedback. A model package can be statistically better and still fail because its live input path is wrong.
Start with shadow scoring when the task permits it. The candidate receives production-shaped inputs without controlling the user action. Compare feature validity, score distributions, error rates, and latency. Shadow scoring cannot directly measure the user outcomes caused by the candidate because it does not change those actions. It is a technical compatibility check, not a substitute for an experiment.
A limited rollout exposes a defined population under an explicit policy. Record assignment and model version so outcomes can be attributed correctly. Avoid switching a user between policies when that creates interference or confusing experiences. Define a fallback that is actually available: an older model, a simple rule, or a reviewed manual path. The fallback also needs the current feature contract.
Choose stop conditions before launch. Examples include schema failures, a latency threshold, an excessive missing-feature rate, or a protected operational metric. Delayed labels require an observation window. An immediate lack of complaints does not establish model quality. Separate fast service checks from slower outcome checks.
Worked example
A fictional ranking model has a p95 scoring budget of 80 ms and a missing-feature ceiling of 1%. Shadow traffic shows 60 ms p95 and 0.4% missing features. The candidate passes this compatibility stage. A later limited rollout assigns 5% of users to the candidate, keeps assignment stable, and records the ranking policy with each impression.
After a feature-service change, missing features rise to 8% for the candidate. The release rule sends affected requests to the prior compatible model and pauses expansion. The team investigates feature version mismatches. It does not wait for click labels to prove that an input contract already failed. The rollback event retains timestamps and assignment information so later outcome analysis can account for the exposure.
Exercise and solution
A candidate model improves offline recall but requires a feature absent on 15% of live requests. The old model does not require it. Propose the next step.
A strong answer validates the missingness pattern, defines a documented missing-input route, and tests it before exposure. It can continue shadow evaluation while blocking rollout. It does not fill every absence with zero unless that behavior was trained and evaluated. Award one point each for identifying the serving gap, a valid fallback, separate fast and delayed checks, and a predeclared rollout criterion. The candidate's offline gain cannot override a broken input contract.
Lab artifact: a release state record
A useful release record connects evidence to a permitted next step. It does not treat “deployed” as a single binary event.
State
Evidence required
User action owner
Exit condition
Offline candidate
Valid data and fixed policy evaluation
Prior model
Approve shadow test
Shadow
Feature validity, latency, resource use
Prior model
Pass service contract
Limited exposure
Stable assignment and exposure logs
Assigned policy
Mature outcome review
Expansion paused
Trigger and scope recorded
Tested fallback
Repair and revalidate
Full exposure
Agreed outcome and service evidence
Candidate
Continuous checks
The table is an original teaching contract. Google's Rules of ML supports attention to infrastructure, measurement, and production behavior; it does not prescribe these exact states or numerical limits. Each transition needs an owner. The person who diagnoses a feature failure may differ from the person who approves broader exposure. A timestamp alone cannot tell an analyst which policy a request actually used, because fallback can override intended assignment.
Store assignment and actual exposure separately:
code
1request_id: r-2042user_bucket: candidate3assigned_model: risk-v74actual_model: risk-v65fallback_reason: missing_required_feature6feature_contract: account-v37decision_time: 10:05:128outcome_window_end: next day 10:05:12
An intention-to-treat analysis compares assigned groups, preserving the original randomized comparison when its assumptions hold. An analysis by actual model may be biased because fallback selects unusual requests. For example, users with missing account history may differ in risk. Both views can help operations, but they answer different questions. Do not remove failed candidate requests from the assigned treatment group merely to improve the apparent treatment result.
A second failure case: fallback shares the failed dependency
The prior model may be an invalid escape route if it uses the same broken feature transformation. Suppose both models read a category embedding service, and its lookup table changes format. Routing to the prior model still produces wrong inputs. The rollback drill must test the entire path, including feature versions and request schema, rather than only loading an old weight file.
A fallback can instead be a compatible package with its required preprocessing, an approved simple rule, or a manual review route within available capacity. Each choice trades quality, coverage, latency, and cost. For a review route, calculate whether an incident sending all traffic to humans would exceed capacity. A nominal fallback that cannot process the affected volume is not an operational escape route.
Exercise: decide what the evidence permits
A candidate has 70 ms p95 latency under a budget of 80 ms. It produces one percent scoring errors; the agreed ceiling is 0.2 percent. Its conversion rate appears higher after two hours, but the defined purchase window is seven days. The prior model passes the error ceiling and handles all current request types.
Pause candidate exposure and use the tested prior path. Latency passes, but the error rate fails independently. The purchase labels are immature, so their apparent improvement cannot override the immediate failure. Record assigned and actual exposure for affected users and preserve error samples. A repair should pass a targeted reproduction and the same service checks before exposure resumes. Award one point for each independent check, one for label maturity, and two for the fallback and audit plan.
Misconceptions to correct
“Shadow traffic proves the candidate has no feedback risk” fails because shadow predictions do not cause actions; future behavior under those actions is unobserved. “Rollback means replacing the model file” fails when parsing, feature retrieval, calibration, or policy configuration changed with the model. Package identity must include those components.
For an interview, explain which metrics you would inspect immediately and which require time. Schema errors, missingness, and latency can support an immediate stop. Label-based precision requires complete observation under the specified outcome window. A successful first hour is evidence about that hour's traffic and service path. It does not establish long-term quality, rare-event reliability, or a stable feedback loop. This limited conclusion is stronger than an unsupported claim of safety because another engineer can reproduce and audit it.
Interview probe
Original practice: What does shadow deployment prove? A strong answer describes compatibility and performance evidence while excluding causal user-outcome claims. Follow up with delayed labels. A weak answer calls shadow scoring a complete A/B test.
What can shadow scoring support when candidate decisions never reach users?
AA conversion lift caused by the candidate.BA stable long-term response to candidate actions.CProduction-shaped serving latency and input compatibility.DA causal comparison of candidate and prior recommendations.
A candidate-assigned request uses fallback because its feature is absent. What should be logged?
AOnly assigned model, to keep the experiment table simple.BOnly actual model, because assignment had no effect.CNeither until the outcome matures.DBoth assignment and actual model, plus fallback reason.
Both candidate and old model depend on the same broken feature lookup. What makes rollback valid?
AConfirm a complete compatible fallback path, including features.BLoad old weights while retaining the failed lookup.CTreat all missing features as zero without evaluation.DExpand exposure to collect more outcome labels first.
Latency passes, scoring errors fail the predeclared ceiling, and outcome labels are immature. Which action follows the contract?
AAverage the passing and failing service metrics.BPause exposure and use the tested fallback while repairing errors.CContinue until purchase labels mature because outcome metrics dominate.DRemove errored requests before comparing assigned groups.
Why can grouping outcomes by actual model create bias after fallback?
AThe original random assignment guarantees that fallback requests remain a random sample.BRemoving fallback requests restores the originally assigned comparison.CSelected requests may differ systematically from normal requests.DRandom assignment automatically balances every actual-exposure subgroup.
Can you design an escape route that includes feature dependencies and preserve assignment as well as exposure? Rate confidence from 1 to 5 and identify the untested component.
Not yetGetting thereConfident
Wrap-up
Connect rollout to a tested fallback. Use immediate service checks and delayed outcome evidence for their separate purposes.