← All questions
MediumAI ML
Why divide by sqrt(d_k) in scaled dot-product attention?
Asked atOMOpenAI MLE
1Give yourself 5 minutes
2Answer out loud, not in your head
3Then compare with the answer below
Reference answer
Then expect these follow-ups
Why multi-head?
How does this interact with FlashAttention?
Why does softmax(QK^T) without scaling still work sometimes in inference?
Free to read · better with Enzo
Practice this out loud with Enzo
Enzo runs it as a mock interview, pushes back with follow-ups, and grades you on the rubric.
Next question