← All questions
MediumAI MLSystem design

Eight GPUs are available on four hosts, two per host. The model requires two GPUs per replica. A host-local replica measures40 requests/minute; an eight-GPU group measures110. Demand is 100/minute after one host loss. Compare the two layouts under the assumption that a parallel replica needs every worker.

1Give yourself 5 minutes
2Answer out loud, not in your head
3Then compare with the answer below
0

Reference answer

Then expect these follow-ups

  • What changes when tensor communication crosses a slower inter-node link?

Free to read · better with Enzo

Practice this out loud with Enzo

Enzo runs it as a mock interview, pushes back with follow-ups, and grades you on the rubric.

Next question