Lessons
1Tokenization deep dive46 min read
The model never sees text — it sees token ids from a byte-level BPE vocabulary. How merges are learned, why 50,257 is not arbitrary, why “strawberry” and arithmetic break at the tokenizer (not the network), the 4:1 chars-to-tokens economics, and the tokenizer-level failure modes that no fine-tune can fix.
- →Tokens and Generation
- →Transformer Architecture
2Attention from scratch48 min read
Scaled dot-product attention, derived. Why Q/K/V are three projections of the same input, why we divide by √d_k (the exact variance argument), what multi-head actually is (a shape split, not extra compute), the exact tensor shapes through one head, and the quadratic cost that defines everything downstream.
- →Transformer Architecture
- →Attention and Training Stability
3Embeddings & positional encoding47 min read
What the embedding vector actually encodes, how encoders pool tokens into one vector, and the position problem: attention is permutation-invariant, so order must be injected. Sinusoidal vs learned-absolute vs RoPE — the rotation identity that makes RoPE relative — and how the positional scheme caps your context length.
- →Embeddings and Position
- →Transformer Architecture
4The transformer block & stacking47 min read
How attention and the FFN are wired into a block, and why the wiring is the part that determines whether a 70B model trains at all. Residual connections, LayerNorm vs RMSNorm, pre- vs post-norm (the gradient-flow argument), the 4× FFN and SwiGLU, and why stacking depth works.
- →Transformer Architecture
- →Attention and Training Stability
5Inference lifecycle: prefill, decode & the KV cache48 min read
What actually happens when you call the model: the autoregressive loop, the prefill/decode split and why their costs are inverted, KV-cache mechanics and the exact memory formula, why decode is bandwidth-bound, and how temperature/top-k/top-p shape the output — with real per-token byte counts.
- →Efficient Transformer Serving
- →Tokens and Generation
6Attention & serving at scale49 min read
The production techniques that make transformers servable: MQA/GQA (cut KV bytes at the source), FlashAttention 1/2/3 (IO-aware exact attention), sliding-window/sparse attention, PagedAttention + continuous batching, and quantization. Real measured impact, the tradeoffs, and how vLLM/TensorRT-LLM/SGLang differ.
- →Efficient Transformer Serving
- →Transformer Architecture
7Capstone: annotate a forward pass46 min read
Assemble the whole track into one exercise: trace tensor shapes and cost through a mini-GPT from token ids to logits, then size a real serving deployment — KV-cache memory, what fits on which GPU, and the levers to hit a latency/throughput target. The canonical LLM-engineer internals interview, rehearsed end to end.
- →Transformer Architecture
- →Efficient Transformer Serving
Skills in this course
- 01Transformer ArchitectureTrace token, attention, block, and output shapes through a transformer.
- 02Tokens and GenerationExplain tokenization, autoregressive generation, and sampling behavior.
- 03Attention and Training StabilityReason about scaled attention, residual paths, normalization, and feed-forward blocks.
- 04Embeddings and PositionExplain token embeddings, pooling, and positional encoding tradeoffs.
- 05Efficient Transformer ServingSize KV cache and choose serving techniques for latency and throughput targets.