← All questions
MediumAI MLSystem design

The same mixed chat/document trace produces 900 output tokens/s with chat p95 gaps 35 ms. A larger batch produces 1,080 tokens/s but chat gaps 70 ms. The chat limit is 50 ms. Decide whether to accept and propose a controlled next test.

1Give yourself 5 minutes
2Answer out loud, not in your head
3Then compare with the answer below
0

Reference answer

Then expect these follow-ups

  • How would you test a mix of short chat messages and long document prompts?

Free to read · better with Enzo

Practice this out loud with Enzo

Enzo runs it as a mock interview, pushes back with follow-ups, and grades you on the rubric.

Next question