← All questions
MediumCodingSystem design

Implement scaled dot-product self-attention

1Give yourself 5 minutes
2Answer out loud, not in your head
3Then compare with the answer below

Reference answer

Then expect these follow-ups

  • Extend to batched multi-head (see the next question). Time and memory complexity?

  • O(n²·d) time, O(n²) memory for the score matrix — this is the lead-in to the FlashAttention question below. What's the difference between padding masks and causal masks, and how do they combine?

  • Why multiple heads instead of one big head?

Free to read · better with Enzo

Practice this out loud with Enzo

Enzo runs it as a mock interview, pushes back with follow-ups, and grades you on the rubric.

Next question