Lesson 4 of 4 · 25 min

Release a model without losing the escape route

Specify a model rollout and rollback contract.

Mechanism and reasoning

A model release changes more than a weight file. Tokenizer configuration, serving code, quantization, prompt defaults and adapter selection can change outputs or performance. Record the entire serving identity so a rollback restores a known combination. A mutable model alias alone is insufficient evidence of what processed a request.
Separate service health from answer quality. A new model can return successful HTTP responses while producing invalid structured output or lower-quality answers. Define a small deterministic contract suite, representative quality evaluation and workload test. The course does not require a new evaluation product. The key is to state what evidence blocks promotion.
A canary sends a controlled fraction of traffic to the new version. Compare like-for-like traffic classes, because random small samples can contain different prompt mixes. Protect users with a clear abort rule for severe failures and a minimum evidence window for subtle changes. Shadowing can collect outputs without showing them to users, but it still incurs cost and may involve sensitive input. Apply the same access rules.
Rollback must account for state compatibility. A cache from one model version should not be assumed valid for another. In-flight streams may need to complete on their original version. Retaining the old warm pool can make rollback fast, at a temporary cost. Record this cost before the release.
An interview answer becomes useful when it names a stop threshold, who acts and how to verify recovery. 'We use blue-green deployment' is a label. Explain the request routing, version identity, state boundary and evidence required to switch back.

Write a release record that supports a real rollback

Treat serving identity as an immutable bundle with separately observable components. The record should make it possible to reconstruct what handled a request without logging the private prompt. Include the model and tokenizer revisions, runtime build, precision settings, adapter selection and application prompt-contract revision. A cache namespace or equivalent compatibility boundary must follow the actual execution identity.
code
1Illustrative release record R242model: immutable revision m243tokenizer: revision t84runtime: build r195precision: explicitly pinned configuration6adapter: a37prompt contract: p128cache compatibility boundary: R249previous stable bundle: R2310abort owner: serving on-call11rollback action: route new eligible work to warm R23 pool
A rollout plan also defines eligible traffic. Suppose the new version supports text but a small group of requests uses another input mode. Excluding that group from the canary can be valid, but passing the canary then establishes only the eligible scope. Do not promote the new version to all traffic without testing the excluded path. Similarly, tenant routing can produce a biased canary if the selected tenant has much shorter prompts than the rest of the service.
CheckEvidencePromotion condition
Service availabilityErrors and readiness by versionWithin stated limit
Output contractParsed schema and required fieldsWithin stated failure limit
QualityRepresentative reviewed casesMeets defined evaluation rule
PerformanceLength-matched latency by classMeets class objectives
RecoveryRehearsed route switch and drainOld bundle resumes expected service
These rows answer different questions. Readiness checks whether the process can serve; it does not prove answer validity. A schema check can prove that a field has the expected type; it does not prove that the value is correct. A small quality evaluation can reveal regressions on its cases but cannot guarantee all future answers. Name the boundary of each check so that one green result does not replace the others.
The threshold denominator needs precision. If 600 canary attempts include 580 completed responses and 20 transport failures, “schema failures per completed response” differs from “unusable results per attempted request.” Both can be useful. Report them separately and ensure the promotion rule does not erase transport failures. Exclude a request only under a documented rule that is applied to both versions.
For a severe defect, use an immediate stop rule that does not wait for statistical confidence. For a small quality or performance shift, use a planned sample window and representative slices. Repeatedly checking a tiny sample and promoting at the first favorable moment encourages a biased decision. This lesson does not prescribe a universal statistical test; it requires the release owner to define the evidence before observing the result.
Rollback has a capacity requirement. If the old warm pool has only one replica while the full service needs three, switching routing back does not restore the previous capacity. Either retain enough old capacity for the rollback contract or document an admission policy during recovery. The temporary overlap of old and new pools should be included in the release cost plan. A rollback command that succeeds administratively can still leave users waiting.
Streams need explicit handling. A request already generating on R24 can finish there if the defect and policy permit. Moving its attention state to R23 is unsafe unless compatibility is specifically supported and verified. If the defect makes continuation unacceptable, cancel the stream and return the defined error rather than relabelling it as an R23 response. Keep the original execution identity in metrics so the recovery graph remains interpretable.
After rollback, verify the signal that caused it and the overall service. If schema failures fall but queue age rises because the old pool is too small, recovery is incomplete. Record the last affected request, routing switch time, in-flight drain outcome and stable-version metrics. These details make the release reversible in operational terms, not merely in deployment configuration.

Worked example

Teaching release record: old version M1 stays warm with two replicas. New M2 gets one replica and 10% of eligible requests. Abort if structured-output failures exceed 1% over at least 500 requests, or if a severe authorization defect appears at any sample size. At 600 requests, M2 has twelve format failures, giving 2%. Stop new routing to M2, let safe in-flight work drain, restore M1 routing and check format failure rate by version. Do not mix M2 cache state into M1.

Exercise

A canary has eight failures in 800 responses, exactly a 1% threshold defined as 'exceeds 1%'. No severe defect is known. State the rule result and one reason not to declare broad success yet.

Model solution and rubric

The rate equals the threshold, so this rule alone does not trigger an abort. It also does not prove quality across all request types. Review traffic slices, latency, output validity and sample uncertainty before promotion. A severe safety or authorization defect would use the separate immediate-stop rule. Write thresholds precisely so two operators do not reach different actions from the same data.
Score out of four: one point for the correct result, one for showing the intermediate reasoning, one for identifying the stated failure case, and one for a verification that could disprove the answer. Do not award the reasoning point for a tool name alone.

Failure modes and misconceptions

“Switching an alias restores the old system.” The alias can leave runtime, tokenizer, cache or prompt settings incompatible. Restore the complete known bundle.
“A green schema check proves answer quality.” Structurally valid output can still be wrong. Use separate contract and quality evidence.

Interview probe

Evidence class: recommended. Original practice.
What must a rollback restore besides model weights?
Strong answer: The compatible runtime, tokenizer, quantization and prompt configuration, plus routing and cache separation. I would preserve version identity on requests, drain compatible streams and verify the specific metric that triggered rollback.
Follow-up: How do you prevent a slow quality regression from hiding behind healthy HTTP status codes?
Weak answer indicators: Rollback by alias alone; testing only availability; sharing incompatible cached state; an abort rule with no sample or severity definition.

Sources

Technical references: vLLM prefix caching; vLLM metrics design; Kubernetes Deployments. Sources support the documented mechanisms. The numbers, decisions, rubrics and interview prompts in this lesson are original teaching examples, not measurements or employer question claims.
docsvLLM prefix cachingdocs.vllm.aidocsvLLM metrics designdocs.vllm.aidocsKubernetes Deploymentskubernetes.io

Checkpoint

A canary returns valid JSON with incorrect factual values. Which evidence is missing?

AA stricter JSON parser that rejects malformed outputBA lower latency threshold for completed answersCA larger count of successfully parsed canary answersDRepresentative answer-quality evaluation
Sign up free to answer and see why

Checkpoint

A rollback alias points to old weights but uses the new tokenizer. What is wrong?

AThe complete known serving identity was not restoredBRollback must always restart every clientCOld weights cannot be used after any canaryDTokenizer revisions never matter
Sign up free to answer and see why

Checkpoint

Ten schema failures occur in 500 completed canary responses. Rule: abort above 1% after at least 500. Result?

AContinue because fewer than twenty failedBAbort because 2% exceeds the ruleCPromote because 98% are validDWait for all service traffic to use the canary
Sign up free to answer and see why

Checkpoint

A severe authorization defect appears after twenty requests. A separate immediate-stop rule exists. What follows?

AWait for the routine 500-request sampleBIgnore it if latency is healthyCStop under the severe-defect ruleDAverage it with old-version successes
Sign up free to answer and see why

Checkpoint

The old pool can serve only half current demand. What does a successful routing rollback establish?

AFull recovery of user latencyBA proof that the new release caused all errorsCThat every old cache entry is warmDRouting changed; capacity and recovery still need verification
Sign up free to answer and see why

Explain how you would specify a model rollout and rollback contract without reading the solution. State one assumption that could change your answer, and one observation that would make you revise it.

Not yetGetting thereConfident

Wrap-up

  • Release a complete serving identity. Keep the old version ready until the new version has passed its stated evidence checks.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.