Lesson 1 of 4 · 25 min

Read desired state as a control loop

Trace reconciliation and distinguish desired from observed state.

Mechanism and reasoning

A declarative API records a desired state. A controller observes the system and acts to reduce the difference. This is a repeated process, not a synchronous command that guarantees the requested result before the API responds. A successful update means the request was accepted under the API's rules. It does not mean the application is serving traffic.
A Deployment manages ReplicaSets, which manage Pods. Those layers have different identities and responsibilities. The Deployment describes rollout intent. A ReplicaSet maintains a replica count for a template. A Pod is an execution unit that can disappear and be replaced. Treating a Pod name as a durable service identity creates fragile operations.
Reconciliation must tolerate retries. A controller can observe stale information or repeat an action after a failure. Its operations should converge rather than create a new resource on every pass. Ownership references and labels help the system connect resources, while resource versions support conflict detection on updates. The detailed API behavior matters when multiple writers change the same object.
Observed state can lag desired state for legitimate reasons. Scheduling may be blocked by resources, image pulling may fail, or readiness may not pass. Diagnose the condition at the layer that owns it. Increasing replicas cannot fix an invalid image reference. Repeatedly deleting Pods can erase useful evidence while the controller recreates the same failure.
In an interview, narrate the state transitions. Say what was accepted, what was created, what is ready and what is serving. This vocabulary prevents a false conclusion such as 'the deployment succeeded because the API returned 200.'

Control-loop case study

Teaching input: an application has a Deployment with desired replicas three. The current ReplicaSet has three Pods. Two are ready and one is waiting for an image that does not exist. An operator edits the replica count to four. The API accepts the update and stores a new resource version. The Deployment controller observes the change and the ReplicaSet creates a fourth Pod. That Pod uses the same invalid image reference and also fails to start.
The final state can be desired four, created four, ready two. Nothing in this trace requires a controller bug. The controller successfully reconciled the count while the workload remained unhealthy. The next useful observation is the Pod's image-pull status and the Deployment conditions. A higher desired count only adds another failed attempt.
Now suppose a second operator applies an older manifest that specifies three replicas and the previous valid image. Two fields changed, but the intended fix may have been only the image. This illustrates why ownership of fields matters. A deployment tool, an autoscaler and a person can all write to related state. Without a clear ownership contract, each can undo the others. Review the actual update strategy and field management rather than assuming the last visible command expresses everyone's intent.
For a custom controller, write down a stable resource identity. Imagine a controller that creates a helper ConfigMap whenever it sees a Deployment. If it generates a random name each time, every reconciliation adds another ConfigMap. A deterministic name based on the owning object and a check for existing state makes retries converge. However, a deterministic name alone does not solve update conflicts or cleanup. The controller must compare the required content, handle already-exists responses and decide what happens when the owner is deleted.
A useful controller test does not merely run the happy path once. Run reconciliation twice with unchanged input and verify that the second pass creates no additional resources. Then simulate an API conflict and rerun. Finally, remove one managed resource and verify that the controller repairs it without changing unrelated objects. These are tests of convergence and ownership, not tests that mirror implementation lines.

Decision artifact

ObservationInterpretationNext evidence
API update acceptedDesired state was storedObject generation and controller conditions
Desired four, created four, ready twoCount converged but health did notPod waiting reasons and readiness results
New Pods repeat image failureTemplate defect is plausibleExact image reference and registry result
Replica count keeps revertingMore than one writer may own intentField management and autoscaler configuration
Each row narrows the question. It does not declare a root cause from a status word. If the image exists and credentials are correct, the next investigation can move to registry reachability. If the controller never observes the new generation, investigate controller health and API connectivity. Preserve this branch in the explanation so a learner can adapt when the first hypothesis is wrong.
The operational objective is a stable service with the declared number of useful replicas. Count convergence is one part of that objective. Availability, correct image identity and request behavior still need checks. A control loop can work exactly as designed while faithfully maintaining a bad specification.

Worked example

Invented timeline:
TimeDesiredCreatedReadyEvent
10:00333Version A healthy
10:01343Rollout creates one surge Pod
10:02342One old Pod fails independently
10:03342New image cannot start
The system has four Pod objects but only two useful replicas. Inspect the failed old Pod separately from the new image failure. Rollout status and serving capacity answer different questions. A recovery plan must restore healthy capacity without removing the remaining two healthy instances.

Exercise

A controller manages one ConfigMap named settings-a for Deployment a. On each retry it creates settings-a-new with a random suffix. After four reconciliations, four extra ConfigMaps exist. Describe the defect, the expected second-run behavior and one conflict test.

Model solution and rubric

The create operation lacks stable identity and convergence. The controller should identify its one managed ConfigMap deterministically, compare required content and update only when needed. A second run with unchanged desired and observed state should make no additional resource. Simulate an update conflict by changing the resource version between read and write, then verify the controller rereads and retries without overwriting an unrelated field. Cleanup must follow ownership policy rather than deleting every similarly named object.
Score out of four: one point for the correct result, one for showing the intermediate reasoning, one for identifying the stated failure case, and one for a verification that could disprove the answer. Do not award the reasoning point for a tool name alone.

Failure modes and misconceptions

Misconception 1: API acceptance proves application readiness. The API can store a desired state that the scheduler or runtime cannot realize. Misconception 2: deleting failed Pods fixes a bad template. The controller usually creates replacement Pods from the same template, so the fault returns and the diagnostic history can be lost.

Interview probe

Evidence class: recommended. Original practice.
The Deployment says three replicas but users see failures. What do you inspect?
Strong answer: I separate desired, created, ready and available replicas, inspect conditions and Pod events, then check request behavior. Replica count alone does not prove correct routing or application health.
Follow-up: How would an autoscaler and a deployment tool fight over replica count?
Weak answer indicators: Using API success as rollout success; changing counts to fix image errors; assuming repeated reconciliation runs only once.

Sources

Technical references: Kubernetes Deployments; Kubernetes probes. Sources support the documented mechanisms. The numbers, decisions, rubrics and interview prompts in this lesson are original teaching examples, not measurements or employer question claims.
docsKubernetes Deploymentskubernetes.iodocsKubernetes probeskubernetes.io

Checkpoint

An API update returns success while the new image cannot be pulled. What has succeeded?

AThe desired state updateBThe application rolloutCThe readiness checkDThe user request path
Sign up free to answer and see why

Checkpoint

Desired replicas are five, created Pods five and ready Pods two. Which conclusion is strongest?

AThe controller ignored the replica countBCount converged but workload health did notCAll five Pods can serveDThe scheduler must be unavailable
Sign up free to answer and see why

Checkpoint

A custom controller creates one randomly named helper ConfigMap per reconciliation. Which change makes repeated runs converge to one owned helper?

AIncrease the resync period and retain the random nameBUse a stable owner-derived helper identity, compare state, and handle create conflictsCDelete every ConfigMap with the same label before creating a new oneDRecord the last random name only in the controller process memory
Sign up free to answer and see why

Checkpoint

A controller reads resourceVersion 41. A user updates an unrelated annotation to version 42 before the controller writes its desired setting. The API returns a conflict. What preserves both intentions?

AForce the complete version 41 object over the current objectBDelete the object and recreate it from the desired templateCReread version 42, apply only the owned setting, and retry with conflict handlingDTreat the conflict as success because another writer has updated the object
Sign up free to answer and see why

Checkpoint

A new requirement is one useful replica after one of two Pods fails. Which observation validates it?

ABoth Pod objects existBDesired replicas equals twoCOne remaining Pod is ready and serves the required requestsDThe API accepted the count
Sign up free to answer and see why

Explain how you would trace reconciliation and distinguish desired from observed state without reading the solution. State one assumption that could change your answer, and one observation that would make you revise it.

Not yetGetting thereConfident

Wrap-up

  • Read the controller's state transitions, then verify useful service. Repeated actions should converge and respect ownership.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.