Lesson 1 of 4 · 25 min

Keep a rollout inside its capacity budget

Calculate rollout capacity and define an abort condition.

Mechanism and reasoning

A rollout replaces working capacity with a new version. The platform must preserve enough useful service while the change proceeds. Kubernetes Deployment settings such as maximum surge and maximum unavailable control counts, but they do not directly know the business workload. Translate those counts into request capacity and include readiness delay.
Surge capacity needs real resources. A policy that allows an extra Pod does not create room for it on a full cluster. The rollout can stall if the new Pod cannot schedule. Conversely, allowing unavailable replicas can reduce service below demand even when the rollout follows its configuration exactly.
Readiness is a necessary signal for normal routing, but a ready application can still return wrong results. Add request-level checks for the changed behavior and monitor the traffic slice on the new version. A rollout should have an abort condition tied to a measurable defect. Avoid thresholds defined only after the results arrive.
Stateful changes can limit rollback. If the new version writes data that the old version cannot read, switching the image back may not restore service. Use an expand-and-contract migration where possible and test both versions during the overlap. A backup is not a substitute for a compatibility plan.
The interview answer should describe normal progression, a stuck new replica and an independent failure during rollout. A plan that tolerates only the expected update is weaker than one that reserves capacity for a realistic concurrent fault. State which failure condition is required rather than promising every combination.

Capacity case

Teaching service needs 120 requests per second. Each ready replica supports forty within the latency objective. Four replicas provide 160, leaving forty of healthy-state margin. A rollout with maximum unavailable one can leave three ready replicas and 120 capacity if everything else behaves as measured. There is no spare capacity for another replica failure at that point.
If the requirement includes one independent replica failure during rollout, keeping at least four ready replicas while changing versions gives a stronger starting condition. That can require surge capacity and readiness before removing old Pods. The exact rollout settings depend on the controller and placement, but the capacity argument is independent of the syntax.
Assume the cluster has room for exactly four replicas and all are occupied. A maximum-surge setting of one asks for a fifth Pod but does not guarantee it can schedule. The rollout may wait indefinitely for space. The operator can add capacity, revise the rollout requirement or schedule a controlled reduced-capacity window. Reducing requests without workload evidence merely hides the shortage from the scheduler.
StageOld readyNew readyTotal capacityDecision
Before change40160 req/sMeets demand with margin
New Pod starting40160 req/sDo not count starting Pod
New Pod validated41200 req/sOne old Pod can drain
One old Pod removed31160 req/sPreserve four useful replicas
One additional failure21120 req/sMeets stated demand with no margin
The table assumes the new version has the same measured capacity. That assumption needs testing. If the new version serves only thirty requests per second, the final row provides 110 and fails the requirement. Version-specific capacity matters when an update changes query cost, memory or concurrency.

Release decision record

A release record can name version digest, configuration version, database compatibility, intended traffic, abort threshold and rollback action. For the teaching service, use a hard stop for a confirmed data-integrity defect at any sample size. For ordinary latency regression, require a defined observation window and comparable traffic. The hard stop and statistical comparison serve different purposes.
Record in-flight behavior. Draining should stop new assignments while allowing existing safe requests to finish within a bound. Long-lived connections may require explicit termination behavior. Removing a Pod immediately can convert a healthy rollout into client errors. Readiness removal, load-balancer propagation and process termination are related but not instantaneous.
A controlled rollout experiment can use an invented trace of one hundred requests with fixed inputs. Compare old and new outputs for contract validity, then test a burst at the target arrival rate. This does not prove all production behavior, but it can reject obvious incompatibility before broad exposure. Keep the old version's artifact and compatible configuration available until the promotion criteria pass.
When an interviewer changes the requirement to 'zero failed requests during a zone outage and rollout,' do not reuse the same four-replica answer. Recalculate from zone placement and surviving capacity. Clarify whether the budget can support that requirement. A credible design makes the cost of the promise visible.
Before promotion, record the measurement interval and the traffic class used to establish forty requests per second per replica. A benchmark using cached reads cannot justify the same capacity for expensive writes. If the release adds a database query, repeat the workload test with representative data. Check the downstream connection budget too: surge replicas can create more database connections while old replicas still run. A rollout can preserve application Pod count and still overload a shared database. Include the peak old-plus-new connection count in the release worksheet, and bound it with a connection pool or staged startup. This is a separate constraint from CPU and memory placement.

Worked example

Illustrative rollout settings:
yaml
1strategy:2  type: RollingUpdate3  rollingUpdate:4    maxSurge: 15    maxUnavailable: 0
For four replicas, the intent permits a fifth during rollout while avoiding intentional unavailability. This fragment assumes resource room and valid readiness. It does not guarantee zero errors, account for an independent failure or validate output correctness. In the teaching case, retain at least four ready forty-request-per-second replicas to preserve 160 capacity while demand is 120.

Exercise

Demand is ninety requests per second. Each replica provides thirty. There are four ready replicas. During rollout, one old replica is removed before the new one is ready, then another fails. Calculate remaining capacity and propose a safer rollout requirement.

Model solution and rubric

Two ready replicas remain, giving sixty requests per second, below demand by thirty. Require enough ready capacity to survive the specified concurrent failure, and provide surge room so replacement capacity is validated before an old replica leaves. If budget cannot support that, state a reduced availability or maintenance-window contract. Verify the new version's capacity instead of assuming it equals the old one.
Score out of four: one point for the correct result, one for showing the intermediate reasoning, one for identifying the stated failure case, and one for a verification that could disprove the answer. Do not award the reasoning point for a tool name alone.

Failure modes and misconceptions

Misconception 1: maximum surge creates physical capacity. It permits additional Pods but scheduling still needs resources. Misconception 2: rollback means restoring the old image only. Data and configuration can become incompatible, so rollback needs a complete tested version boundary.

Interview probe

Evidence class: recommended. Original practice.
Your rolling update has zero maximum unavailable. Can you promise zero downtime?
Strong answer: No. I still need schedulable surge capacity, correct readiness, request draining, output checks and the required failure tolerance. The count setting alone cannot guarantee user behavior.
Follow-up: What changes if new replicas have lower per-replica throughput?
Weak answer indicators: Counting starting Pods as capacity; ignoring schema compatibility; treating controller settings as service guarantees.

Sources

Technical references: Kubernetes Deployments; Kubernetes container resources; Kubernetes probes; Kubernetes disruptions. Sources support the documented mechanisms. The numbers, decisions, rubrics and interview prompts in this lesson are original teaching examples, not measurements or employer question claims.
docsKubernetes Deploymentskubernetes.iodocsKubernetes container resourceskubernetes.iodocsKubernetes probeskubernetes.iodocsKubernetes disruptionskubernetes.io

Checkpoint

Four replicas each serve forty, but one is starting. Ready capacity is?

A120B160C200D40
Sign up free to answer and see why

Checkpoint

The cluster has room for four Pods, all occupied. A rollout permits one surge Pod and no intentional unavailability. What can happen?

AThe surge Pod stays Pending and rollout waitsBThe scheduler overcommits declared requests automaticallyCThe rollout must delete an old Pod despite maxUnavailable: 0DThe surge setting reserves a new node before the rollout
Sign up free to answer and see why

Checkpoint

R2 writes records that R1 cannot read. Which rollback plan is sufficient?

ARestore only the R1 imageBRestore R1 and relabel R2 records without conversionCUse a tested data-compatibility or conversion plan with the version rollbackDRestart R1 until it accepts the new records
Sign up free to answer and see why

Checkpoint

Demand is 100; three ready replicas provide thirty each. Result?

ATen percent spare capacityBExactly enough capacityCEnough if desired count is fourDTen requests per second of deficit
Sign up free to answer and see why

Checkpoint

A canary has low error volume but one confirmed cross-tenant disclosure. Under a hard security stop rule, what should happen?

AWait for 500 samples to make the rate stableBAverage its errors with the healthy old versionCPromote if latency improvedDStop new canary traffic and use the defined containment/rollback path
Sign up free to answer and see why

Explain how you would calculate rollout capacity and define an abort condition without reading the solution. State one assumption that could change your answer, and one observation that would make you revise it.

Not yetGetting thereConfident

Wrap-up

  • Translate rollout counts into useful capacity and compatible state. Write the abort and recovery conditions before promotion.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.