← All questions
MediumAI MLSystem design

A 24 GiB GPU holds13 GiB of weights and reserves3 GiB for runtime. Standard non-MLA attention has 32 layers,8 KV heads,128 head dimension and 2-byte cache values. Every request may hold8,192 input-plus-output tokens. Calculate the simplified concurrency bound and explain why a test at twelve concurrent requests can fail.

1Give yourself 5 minutes
2Answer out loud, not in your head
3Then compare with the answer below
0

Reference answer

Then expect these follow-ups

  • What changes if the model uses four KV heads rather than eight?

Free to read · better with Enzo

Practice this out loud with Enzo

Enzo runs it as a mock interview, pushes back with follow-ups, and grades you on the rubric.

Next question