Lesson 1 of 4 · 25 min

Separate naming, routing, and application failures

Use layered network evidence to localize a service failure.

Mechanism and reasoning

A platform debugging interview rewards a sequence of discriminating tests. Start with the exact symptom, source identity and destination. 'The network is down' is too broad. A hostname failure, connection refusal, timeout and application error suggest different paths and need different evidence.
Resolve naming separately from connectivity. If a request to a hostname fails but the same endpoint works by address, name resolution or the returned address becomes a useful first hypothesis. This is not proof that DNS servers are down. A wrong namespace, search suffix, stale record or denied DNS egress can produce similar symptoms.
Next distinguish Service routing from Pod reachability. A Service can have no ready endpoints, a wrong selector or a wrong target port. A direct Pod connection can succeed while the Service path fails. Conversely, a direct connection that bypasses normal routing can omit policy or identity behavior, so do not generalize it beyond what the test establishes.
Use observations that change one variable at a time. Record the source Pod, namespace, destination, protocol and whether the connection is new. Compare from an intended client and an unrelated client when policy is involved. Avoid changing selectors, policies and ports together before you understand the trace.
The interview answer should produce a short fault tree and a verification plan. It should also explain when to mitigate. If users are affected, restore a known valid configuration or route where evidence supports it, while preserving the state needed to identify the cause.

Trace packet

Teaching incident begins after a Service configuration change. The gateway cannot reach billing.default.svc by name. A debug client in the same namespace resolves that name to the expected Service address. Connecting to the Service on port eighty times out. Connecting directly to a ready billing Pod on port 8080 succeeds. The Service specification maps port eighty to targetPort 8000, while the container listens on 8080.
This evidence points to a port-mapping defect. DNS resolution succeeds, so restarting DNS is not supported. Direct Pod success shows that the process can answer on 8080 from this source, though it does not prove every client is authorized or every endpoint is healthy. The Service's target port is the first concrete mismatch to repair.
TestResultInference
Resolve Service nameExpected Service addressNaming works for this client at this time
Connect Service port 80TimeoutFailure remains on routed path
Connect ready Pod port 8080SuccessApplication listens and answers on 8080
Inspect Service targetPort8000Routing target differs from listener
Inspect ready endpointsPresentEmpty endpoint set is not the current explanation
After correcting the target port through the normal reviewed path, repeat the Service request from the original gateway. A debug client alone is not enough because NetworkPolicy or identity can differ. Verify new connections and representative requests. Record whether error rates return to baseline and whether any clients cache old data.

Alternate branch

Now change one fact: direct Pod access also times out. The target-port mismatch may still exist, but it no longer explains every failure. Check source egress and destination ingress rules, application listener binding and node reachability. If the process listens only on loopback, a local health check may pass while remote Pod access fails. If a policy permits only gateway-labelled Pods, the debug client may be intentionally denied.
A useful fault tree starts with 'Can this client resolve the intended name?' If no, inspect resolver configuration, search path and DNS access. If yes, ask whether the returned destination is correct. Then test connection establishment and application response. A successful TCP connection followed by HTTP 403 is not a transport outage; it points to application or gateway authorization. A connection refusal can indicate no listener or an active reject, while a timeout can have several causes. Treat these as clues, not universal diagnoses.
The operator should not copy a long list of commands into the interview and hope one works. Explain the expected result of each test and what branch it chooses. This makes the reasoning portable across tooling. It also reduces accidental changes during diagnosis.

Evidence quality

Use the same request method and path where practical. A root health endpoint can succeed while a business route fails because it needs a database. A direct-IP request can change the Host header or TLS server name and therefore test a different virtual host. When making that comparison, preserve the relevant host and TLS identity or state that the test is only transport-level.
The final answer names the smallest supported fault, the repair and the verification. In the teaching trace, the port mismatch is observed directly. It is appropriate to call it a defect. Whether it caused every production error still requires the original client and user-request checks. Keep that distinction even when the first repair appears obvious.

Worked example

Configuration comparison:
yaml
1# Service fragment observed in the teaching incident2ports:3  - port: 804    targetPort: 80005# Application listener observed in the same incident6# TCP 0.0.0.0:8080
The configured destination port differs from the listener. The intended correction is targetPort 8080, subject to the service's actual contract. After the correction, repeat the original gateway-to-Service request. Do not infer success from a manifest diff alone.

Exercise

A Service name resolves correctly. The Service has zero ready endpoints. Two Pods exist, but their readiness checks fail. What should you inspect before changing DNS or the Service selector?

Model solution and rubric

Inspect the readiness failure and whether the Pods match the intended selector. Existing Pods are not necessarily ready endpoints. Check whether startup is incomplete, the readiness path is wrong or a required local condition fails. If labels already match, changing the selector can route traffic to unsuitable Pods. Verify readiness and a real request after repair. The DNS result already shows naming works for the tested client.
Score out of four: one point for the correct result, one for showing the intermediate reasoning, one for identifying the stated failure case, and one for a verification that could disprove the answer. Do not award the reasoning point for a tool name alone.

Failure modes and misconceptions

Misconception 1: a timeout proves DNS failure. Resolution can succeed while routing or application reachability fails. Misconception 2: a successful direct-IP request proves the full hostname path. Host headers, TLS identity, policy and load balancing can differ, so preserve the relevant conditions.

Interview probe

Evidence class: recommended. Original practice.
A service works by Pod address but fails through its Service. How do you proceed?
Strong answer: I compare selectors, ready endpoints, Service port and target port, then test the original client path. I keep DNS, routing and application response separate and avoid changing unrelated layers.
Follow-up: What changes if the direct-IP request omitted the original TLS server name?
Weak answer indicators: Restarting DNS without resolution evidence; treating Pod existence as readiness; ignoring virtual-host differences.

Sources

Technical references: Kubernetes Services; Kubernetes NetworkPolicy; Kubernetes probes; Kubernetes Deployments. Sources support the documented mechanisms. The numbers, decisions, rubrics and interview prompts in this lesson are original teaching examples, not measurements or employer question claims.
docsKubernetes Serviceskubernetes.iodocsKubernetes NetworkPolicykubernetes.iodocsKubernetes probeskubernetes.iodocsKubernetes Deploymentskubernetes.io

Checkpoint

From one client, DNS returns the expected Service IP. The Service request times out, but the same application route succeeds on a ready Pod at port 8080. What should you compare first?

AService targetPort and the ready endpoint addresses and portsBDNS search suffixes before examining the resolved destinationCApplication database credentials before examining the routed pathDPod memory limits before examining the Service configuration
Sign up free to answer and see why

Checkpoint

A Service has no ready endpoints. Its selector matches two existing Pods, both unready. What is the next useful inspection?

AReplace the selector with a broader application labelBChange targetPort to the readiness probe portCRead the readiness results and the startup/dependency state they measureDEnable publishing unready addresses to restore normal traffic immediately
Sign up free to answer and see why

Checkpoint

A hostname HTTPS request fails. A direct-IP test succeeds only with certificate checks disabled and a different Host header. What has the test established?

AThe hostname route is healthy because the IP is reachableBDNS is the only possible faultCThe Service certificate is valid for the hostnameDOne altered request can reach a responder; repeat with the intended TLS and HTTP identity
Sign up free to answer and see why

Checkpoint

A debug Pod succeeds after a target-port repair. The original gateway still times out. Which comparison is most useful next?

ASource-specific egress, destination ingress and gateway request identityBA further port change inferred from the successful debug testCA DNS cache flush on every node before checking the gateway resolutionDA scale-out based only on the fact that two clients differ
Sign up free to answer and see why

Checkpoint

After correcting a directly observed target-port mismatch, which evidence best supports closing that defect as the cause of the original user failure?

AThe reviewed manifest now contains the intended numberBThe original gateway completes representative Service requests and its error rate recoversCThe debug Pod can still call one Pod addressDThe Service has the same number of ready endpoints as before
Sign up free to answer and see why

Explain how you would use layered network evidence to localize a service failure without reading the solution. State one assumption that could change your answer, and one observation that would make you revise it.

Not yetGetting thereConfident

Wrap-up

  • Change one test condition at a time. Preserve identity and request context when comparing network paths.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.