Plan maintenance without confusing it with failure tolerance
Calculate disruption allowance and surviving service capacity.
Mechanism and reasoning
A disruption policy constrains certain voluntary evictions. It does not make Pods immune to node crashes or every controller action. A PodDisruptionBudget can help a maintenance operation avoid removing too many healthy replicas through the eviction path, but it cannot create replacement capacity or guarantee the application's service objective.
Use absolute counts in the first interview calculation to avoid rounding ambiguity. If a workload has four healthy Pods and minAvailable is three, one voluntary eviction can be allowed under the simplified steady state. If only three are healthy, another eviction would violate that condition. The actual status also accounts for in-flight disruptions and the controller's view.
Maintenance needs a replacement plan. Draining a node can move work only if eligible capacity exists elsewhere. A workload with local storage, restrictive affinity or insufficient resources may not reschedule. The disruption budget can correctly block maintenance because the declared availability condition cannot be maintained.
Involuntary loss is a separate design problem. If two Pods share a node and that node fails, both can disappear regardless of a budget that would have denied two planned evictions. Spread replicas across failure domains and calculate surviving capacity. The budget is one control in that design.
An interview answer should distinguish a blocked drain from a broken cluster. It may be a useful safety signal. Investigate why replacement readiness cannot recover before overriding the policy. State the approval and service-impact requirement for any exceptional maintenance action.
Drain case
Teaching service has five replicas, each handling twenty requests per second. Demand is sixty. Three replicas sit on node A and two on node B. The workload has minAvailable four. A maintenance drain of node A tries to evict its Pods through the normal eviction API.
With all five healthy, the first eviction can reduce healthy count to four under this simplified view. A second eviction should wait until replacement readiness restores the allowance. If node B lacks resources for a replacement and no other node is eligible, maintenance stalls. The policy is preserving four available replicas, even though demand arithmetic suggests three might handle sixty. The declared budget and the business capacity requirement are related but not identical.
Why would a team choose four? It may want one additional replica of reserve, or it may have measured that twenty requests per second is not safe for every workload. Do not lower the budget merely because a simple average suggests three. Review the service requirement and workload evidence.
State
Healthy Pods
Capacity at20 each
Budget implication
Initial
5
100 req/s
One planned eviction may proceed
One removed, replacement pending
4
80 req/s
Next eviction waits
Replacement ready
5
100 req/s
Allowance can recover
Node A fails involuntarily
2
40 req/s
Budget cannot prevent the loss
The last row demonstrates the placement problem. A budget does not repair concentration. For the sixty-request requirement after one node loss, the two-node layout is inadequate. Even if the initial count is five, losing the node with three leaves only forty capacity.
Maintenance decision record
A safe plan states the maintenance unit, maximum concurrent drains, replacement capacity, readiness condition and stop rule. For this service, add eligible capacity and spread replicas before draining a node with three replicas if the requirement demands it. Verify that the scheduler can place the replacement and that the application can use its storage. A capacity graph alone does not prove those constraints.
A controlled test can start with one eviction in a disposable environment. Observe the budget status, replacement scheduling and request latency. Then simulate a failed replacement. The expected result is that further voluntary eviction pauses under the policy, rather than continuing until the service is empty. This is a behavior test of the maintenance workflow.
There is also a rollout distinction. A Deployment's rolling update manages availability through its own strategy; a PodDisruptionBudget does not generally constrain every update action by every controller. Do not assume a budget overrides rollout settings. Read the relevant controller and eviction behavior for the selected version. The interview should name the mechanism being used, such as drain via eviction, rather than saying 'Kubernetes will protect us.'
Fresh scenario
Consider six replicas spread evenly across three nodes, two per node. Each provides twenty and demand is seventy. Losing one node leaves four replicas and eighty capacity, meeting the simplified demand with ten of margin. A planned drain still needs the chosen budget and enough scheduling room to restore the desired count. If the next maintenance step starts before replacements are ready, a second node drain can remove the remaining margin.
The result is a sequence, not a single command. Finish and verify one maintenance unit before proceeding under the chosen concurrency rule. Record the observed ready count and user metric at each step. A drain that eventually completes but causes repeated latency breaches did not meet the service objective.
With four healthy selected Pods and no other relevant in-flight disruption, the simplified allowance is one. With three healthy Pods, it is zero. A complete manifest needs kind, API version and metadata. The budget does not reserve nodes and cannot stop an involuntary node failure.
Exercise
Six replicas are spread two per node across three nodes. Each provides fifteen requests per second. Demand is fifty. Calculate capacity after one node loss and after two node losses. Explain what a disruption budget can and cannot guarantee.
Model solution and rubric
One node loss leaves four replicas and sixty requests per second, meeting demand by ten. Two losses leave two replicas and thirty, below demand by twenty. A budget can constrain supported voluntary evictions according to its policy, but cannot prevent involuntary node loss. Verify placement, replacement resources and workload capacity. Planned maintenance should not silently consume the reserve required for a concurrent fault unless the service contract allows it.
Score out of four: one point for the correct result, one for showing the intermediate reasoning, one for identifying the stated failure case, and one for a verification that could disprove the answer. Do not award the reasoning point for a tool name alone.
Failure modes and misconceptions
Misconception 1: a PodDisruptionBudget guarantees the declared count against any failure. It governs supported voluntary disruption paths, not hardware failure. Misconception 2: a blocked drain is necessarily a bug. It can correctly indicate that replacement readiness or capacity cannot preserve the policy.
Interview probe
Evidence class: recommended. Original practice.
A node drain is blocked by a disruption budget. What do you do?
Strong answer: Inspect healthy count, in-flight disruptions, replacement scheduling and the service requirement. I would add or restore capacity before changing the budget, and distinguish planned eviction from involuntary failure.
Follow-up: Does the same budget necessarily constrain a Deployment rolling update?
Weak answer indicators: Overriding the budget before checking replacement capacity; treating desired replicas as healthy; claiming protection from node crashes.
Four selected Pods are healthy. minAvailable is 3. Assume no other unavailable Pods or in-flight evictions and a current budget status. How many additional voluntary evictions can begin under this simplified budget?
A node crash removes Pods even though a PDB would deny their voluntary eviction. Which conclusion follows?
AThe PDB must have selected the wrong PodsBThe application controller must have ignored an eviction denialCThe PDB cannot prevent a node crash; the lost healthy Pods can reduce later eviction allowanceDThe lost Pods must still count as healthy until a drain starts
During drain, a replacement is Pending with FailedScheduling events that name required node affinity and insufficient requested memory. What should you inspect?
AThe PDB alone, because it assigns replacement nodesBOnly the replacement readiness URLCThe Deployment progress deadline before the scheduling constraintsDEligible nodes, their allocatable resources, existing requests and required affinity
Six replicas are spread two per node. Each surviving replica can sustain 20 requests/s under the stated workload. One node fails. Ignore recovery and shared bottlenecks. What capacity remains?
The first drain command finishes, but replacement Pods remain unready. Before starting a second drain, what evidence is required?
ARestored readiness, current eviction allowance and enough surviving capacity for the next stepBThe desired replica count and the first command exit status aloneCA successful request from one Pod, regardless of the remaining capacityDA PDB manifest with the same minAvailable value as before
Explain how you would calculate disruption allowance and surviving service capacity without reading the solution. State one assumption that could change your answer, and one observation that would make you revise it.
Not yetGetting thereConfident
Wrap-up
Calculate maintenance allowance and failure capacity separately. A blocked eviction can preserve a useful contract.