← All questions
MediumSystem designTechnical
A Deployment desires3 replicas;3 Pods exist, only 1 is ready, and clients see failures through its Service. What evidence do you collect before changing replica count?
1Give yourself 5 minutes
2Answer out loud, not in your head
3Then compare with the answer below
0
Reference answer
Then expect these follow-ups
How would an autoscaler and a deployment tool fight over replica count?
Free to read · better with Enzo
Practice this out loud with Enzo
Enzo runs it as a mock interview, pushes back with follow-ups, and grades you on the rubric.
Next question
- A 24 GiB GPU holds13 GiB of weights and reserves3 GiB for runtime. Standard non-MLA attention has 32 layers,8 KV heads,128 head dimension and 2-byte cache values. Every request may hold8,192 input-plus-output tokens. Calculate the simplified concurrency bound and explain why a test at twelve concurrent requests can fail.
- The same mixed chat/document trace produces 900 output tokens/s with chat p95 gaps 35 ms. A larger batch produces 1,080 tokens/s but chat gaps 70 ms. The chat limit is 50 ms. Decide whether to accept and propose a controlled next test.
- Eight GPUs are available on four hosts, two per host. The model requires two GPUs per replica. A host-local replica measures40 requests/minute; an eight-GPU group measures110. Demand is 100/minute after one host loss. Compare the two layouts under the assumption that a parallel replica needs every worker.
- In matched one-hour measurement runs, configuration A costs 12 units/hour and serves 18,000 completed responses, of which 3,000 are late. Configuration B costs 10 units/hour and serves 12,000 timely responses. No other failures occur. Which is cheaper per thousand timely successes, and what does that comparison omit?
- A shared pool has 12,000 free reserved tokens; tenant T has 4,000 under its active limit. A request needs 6,000 and expires in one second; no release is expected for three seconds. What should admission do, and how should a retry reuse identity?
- Two replicas each serve30 requests/minute. Arrivals rise to 90. A third replica takes three minutes to become ready; nothing expires or is rejected. Calculate backlog at readiness and whether it drains. What if a fourth is ready then?