← All questions
MediumAI MLSystem design

Two replicas each serve30 requests/minute. Arrivals rise to 90. A third replica takes three minutes to become ready; nothing expires or is rejected. Calculate backlog at readiness and whether it drains. What if a fourth is ready then?

1Give yourself 5 minutes
2Answer out loud, not in your head
3Then compare with the answer below
0

Reference answer

Then expect these follow-ups

  • How would you distinguish a tokenizer bottleneck from exhausted GPU capacity?

Free to read · better with Enzo

Practice this out loud with Enzo

Enzo runs it as a mock interview, pushes back with follow-ups, and grades you on the rubric.

Next question