← All questions
HardCodingSystem design

Implement FlashAttention-style tiled attention

Asked atMistral AI
1Give yourself 5 minutes
2Answer out loud, not in your head
3Then compare with the answer below

Reference answer

Then expect these follow-ups

  • Also tile the queries (the real algorithm is a 2D loop) and handle the causal mask per tile. Why is attention memory-bound rather than compute-bound on modern GPUs?

  • Backward pass: why does FlashAttention recompute the attention matrix instead of storing it?

  • How does this interact with KV caching at inference — what's different in decode (one query row) vs prefill?

Free to read · better with Enzo

Practice this out loud with Enzo

Enzo runs it as a mock interview, pushes back with follow-ups, and grades you on the rubric.

Next question