← All questions
HardCodingSystem design
Implement FlashAttention-style tiled attention
Asked at
Mistral AI
1Give yourself 5 minutes
2Answer out loud, not in your head
3Then compare with the answer below
Reference answer
Then expect these follow-ups
Also tile the queries (the real algorithm is a 2D loop) and handle the causal mask per tile. Why is attention memory-bound rather than compute-bound on modern GPUs?
Backward pass: why does FlashAttention recompute the attention matrix instead of storing it?
How does this interact with KV caching at inference — what's different in decode (one query row) vs prefill?
Free to read · better with Enzo
Practice this out loud with Enzo
Enzo runs it as a mock interview, pushes back with follow-ups, and grades you on the rubric.
Next question