Assign startup, readiness and liveness checks to distinct conditions.
Mechanism and reasoning
A probe is an automated decision with consequences. Startup checks protect slow initialization. Readiness controls whether a Pod is considered ready for normal service routing. Liveness can trigger a container restart. A single endpoint copied into all three settings can turn one dependency problem into a larger outage.
Liveness should answer whether restarting this process is a useful response to the detected condition. If a shared database is temporarily unavailable, restarting every application process may add cold starts and reconnect storms while the database is still down. Readiness may need to remove a Pod from traffic if it cannot serve useful requests, but even that needs a service-wide plan when every replica shares the dependency.
Startup behavior is different from steady-state health. A model server or a large application can need time to load data or compile work. A startup probe can defer liveness and readiness evaluation until initialization succeeds under its configured policy. A generous delay alone is less informative because it waits even when startup finishes quickly and can still be too short during a legitimate slow start.
Probe thresholds trade detection delay against false positives. Compute the approximate failure window from period and threshold, while remembering execution duration and scheduling add detail. A very short timeout can fail under brief contention. A very long threshold can leave broken service in place. Choose from observed behavior and the recovery objective.
A good interview answer traces what the platform will do when the probe fails. It names the endpoint's dependencies and explains why the action helps. A successful probe should represent the work the Pod can safely accept, not merely that an HTTP server thread exists.
Shared-dependency failure case
Teaching service has six replicas. Each replica's liveness endpoint performs a database query. The database pauses for twenty seconds during a maintenance event. The liveness check runs every five seconds with a failure threshold of three. All six replicas can cross the threshold and restart during the same dependency interruption. Their restarts do not repair the database. They add connection setup, cache warm-up and startup load just as the database recovers.
Change the design by separating process health from dependency availability. Liveness can check a local event loop or internal deadlock indicator where restart is a meaningful recovery. Readiness can reflect whether this replica can serve the supported request class. If the service can answer cached reads during a database outage, a readiness rule that requires every database query to succeed may unnecessarily remove useful capacity. The probe contract should match actual request behavior.
There is a second failure mode. Suppose readiness fails whenever queue depth is nonzero. During a normal burst, all replicas become unready, traffic is removed, queues drain, and replicas return together. The system can oscillate. Queue thresholds need hysteresis and a definition of useful capacity. A platform team should not use readiness as an unexamined load-shedding mechanism.
The following decision table is a teaching design, not a universal endpoint prescription. The local process checks depend on the application's failure modes. A process can return a trivial health response while its main worker is dead, so the check must observe the component that actually serves work.
Condition
Probe/action candidate
Why
Initialization still loading required state
Startup remains incomplete
Normal slow start should not cause a liveness restart
Local worker is deadlocked and restart repairs it
Liveness failure
The action targets the local fault
Replica lacks data needed for requests
Readiness failure
Stop assigning work it cannot complete
Shared database briefly unavailable
Dependency-aware service policy
Restarting every replica can amplify the incident
Timing worksheet
Assume a startup probe has period ten seconds and failure threshold twelve. Its approximate allowed failure window is 120 seconds, with exact timing affected by when the first probe runs and probe duration. If measured startup p99 is ninety seconds, this may provide some margin, but it is not proof of a bound. A large data change can create a new startup distribution. Record the supported data size and inspect startup metrics after releases.
For liveness, a five-second period and three consecutive failures can detect a persistent fault on the order of fifteen seconds after the relevant sequence begins. Do not promise exactly fifteen seconds from the fault's onset. The fault can occur between probes, and timeouts add detail. In an interview, use approximate language and explain the schedule.
Test probes as failure behavior. Simulate a slow startup, a local worker stall and a shared dependency failure in a disposable environment. Verify restart counts, ready endpoints and user request outcomes. A probe endpoint unit test that returns 200 or 500 does not establish that the deployment reacts safely. The observable result must include the platform action.
Worked example
Illustrative configuration fragment, not a complete deployment:
In this teaching design, each endpoint observes the condition chosen for its action. Distinct URLs are not a Kubernetes requirement: two probes can share an endpoint when the same predicate is appropriate for both actions, with suitable thresholds. This fragment omits timeout tuning, security and the application-specific condition definitions. It illustrates separate decisions rather than a configuration to copy blindly.
Exercise
A service needs seventy seconds to initialize. Its current liveness starts immediately and fails after three checks ten seconds apart. Describe the risk, a startup-probe approach and a test for shared database failure.
Model solution and rubric
The service can restart before initialization completes and never become useful. A startup probe with a measured allowance above seventy seconds can protect legitimate initialization and defer the other probes under Kubernetes behavior. Test a twenty-second database outage while the local worker remains healthy. Liveness should not restart all replicas merely because the database is unavailable. Verify readiness and degraded request behavior under the service's declared contract.
Score out of four: one point for the correct result, one for showing the intermediate reasoning, one for identifying the stated failure case, and one for a verification that could disprove the answer. Do not award the reasoning point for a tool name alone.
Failure modes and misconceptions
Misconception 1: every dependency failure should fail liveness. Restart is useful only when it can repair the observed local condition; shared faults can become restart storms. Misconception 2: readiness means the process exists. A live process may lack the state or capacity needed to serve the declared request class.
Interview probe
Evidence class: recommended. Original practice.
What should happen when the database is down but your process is healthy?
Strong answer: The answer depends on which requests remain useful. I would avoid liveness restarts for a shared dependency outage, apply an explicit readiness or degraded-mode policy and verify that reconnection does not overwhelm the recovering database.
Follow-up: How would a queue-based readiness threshold oscillate during a burst?
Weak answer indicators: One dependency-heavy endpoint for all probes; exact timing claims from period alone; treating probes as passive metrics.
Sources
Technical references: Kubernetes probes; Kubernetes Deployments. Sources support the documented mechanisms. The numbers, decisions, rubrics and interview prompts in this lesson are original teaching examples, not measurements or employer question claims.
An application normally needs 75 seconds to initialize. You want fast failure detection after startup without restarting valid initialization. Which design best separates these requirements?
AUse a startup probe with measured initialization allowance, then a shorter steady-state liveness policyBUse one liveness threshold above 75 seconds for startup and every later faultCMark readiness successful during initialization to stop restartsDKeep immediate strict liveness and add more replicas
During a shared database pause, six otherwise healthy application processes fail their database-backed liveness check and restart. Which explanation fits this design?
ARestarts clear the database fault because the clients created itBLiveness removes traffic while preserving each process and its warm cacheCThe failures prove a local deadlock in all six processesDA dependency fault triggered local recovery actions that can add reconnect and warm-up load
The health HTTP thread answers /live, but a local job worker has stopped making progress. A restart repairs this known worker deadlock. Which check best matches the recovery action?
AKeep checking only the HTTP thread because it proves process existenceBCheck a bounded worker-progress signal with a threshold that allows legitimate long jobsCFail liveness whenever the shared database rejects a queryDFail liveness whenever the job queue is nonempty
Readiness fails whenever queue depth exceeds 0. During bursts all replicas leave service, drain, and return together. Which change should be evaluated first?
AReduce the liveness failure threshold to restart overloaded replicas fasterBIncrease readiness probe frequency without changing its queue thresholdCUse a readiness/load policy tied to usable capacity, with hysteresis and a burst testDReduce CPU requests so more replicas appear lightly utilized
Explain how you would assign startup, readiness and liveness checks to distinct conditions without reading the solution. State one assumption that could change your answer, and one observation that would make you revise it.
Not yetGetting thereConfident
Wrap-up
A probe chooses an action. Make the observed condition justify startup waiting, traffic removal or restart.