Lesson 3 of 7 · 47 min
Embeddings & positional encoding
What the embedding vector actually encodes, how encoders pool tokens into one vector, and the position problem: attention is permutation-invariant, so order must be injected. Sinusoidal vs learned-absolute vs RoPE — the rotation identity that makes RoPE relative — and how the positional scheme caps your context length.
Attention can’t tell “dog bites man” from “man bites dog”
q·k doesn’t know where either token sits. That’s a problem — language is ordered. So every transformer injects position somewhere, and the scheme you pick is not cosmetic: it decides whether the model can run at a longer context than it trained on, with zero extra parameters or a brittle retrain. This lesson covers what embeddings encode and the three positional schemes you’ll be grilled on, ending with why RoPE won.(vocab_size, d_model) table that maps each token id to a learned d_model-vector. There’s nothing magic in a single row at init — meaning emerges from training: tokens that play similar roles get pulled to nearby points, so the geometry of the embedding space comes to encode syntactic and semantic relationships. This is the same object as a word2vec/GloVe embedding, except it’s learned jointly with the rest of the network and is contextualized by later layers (the embedding of a token is the same on the way in; what attention produces downstream is context-dependent).CLS token — so a 3-word query and a 400-word passage land in the same space and are comparable by cosine similarity (the embeddings track from the RAG track lives here). A generative LLM keeps a per-token (B, T, d_model) tensor all the way through and never pools. Interview angle. “What’s the difference between the embedding inside GPT and an embedding from text-embedding-3?” → per-token vs pooled-and-normalized for similarity; same word “embedding,” different job.king − man + woman ≈ queen demonstrates that semantic relations show up as consistent vector offsets, which is why cosine similarity (angle, not magnitude) is the right comparison. Magnitude is mostly a frequency/usage artifact, which is exactly why retrieval normalizes vectors before comparing. Inside an LLM the input embedding is the same vector for a token regardless of context; the contextual meaning (“bank” the riverside vs the financial institution) is built up by the attention layers that follow, not by the embedding row itself. That separation — static lookup, then contextualization — is the mental model to carry.d_model is a capacity knob with a real cost. GPT-2 124M uses d_model=768; GPT-3 175B uses 12,288. Because the embedding table is vocab × d_model and is typically tied to the unembedding (Lesson 1), widening d_model grows two of the largest matrices in the model and every activation tensor downstream. So “just use bigger embeddings” is not free — it scales the whole network’s width. Interview angle. “Why not make d_model huge for richer representations?” → it widens every layer’s matmuls and activations and the tied embedding/unembedding tables; depth and attention design usually buy more than raw width.
RoPE: Understanding Rotary Positional Embeddings in transformersHugging FaceSinusoidal absolute encodings (the original)
PE(pos, 2i) = sin(pos / 10000^(2i/d_model)) and PE(pos, 2i+1) = cos(pos / 10000^(2i/d_model)). Each position 0,1,2,… gets a unique d_model-vector, with frequencies in a geometric progression from short to very long wavelengths. The reason for sinusoids over a plain integer index: any shifted position PE(pos+k) is a linear function of PE(pos), which is the structure a model needs to learn relative offsets. d2l.ai notes that learned absolute embeddings (a trainable (max_len, d_model) table, as GPT-2 uses) perform comparably — but both share a hard limit.max_len. A learned table has literally no row for position 4097 if you trained to 4096; sinusoidal can be evaluated past training length but the model never saw those phase combinations and degrades. So with absolute encodings, extending context means retraining — and that’s the pain RoPE was designed to remove.RoPE — rotate Q and K by position
d dimensions into d/2 pairs; pair i is rotated by angle m·θᵢ at position m, where θᵢ = 10000^(−2(i−1)/d) (the same 10000 base as sinusoids). Concretely it’s a block-diagonal matrix of 2×2 rotations applied to q and k before the dot product — no learned parameters, two multiplies and a sin/cos per pair, near-zero FLOPs.m and a key at position n depends only on the relative offset (m − n): ⟨R_m·q, R_n·k⟩ = ⟨q, R_(n−m)·k⟩. So RoPE encodes position absolutely (each token is rotated by its own index) yet makes attention behave relatively (only distance matters) — with zero extra parameters and zero learned position table. And because high-frequency pairs decorrelate over long distances, RoPE has a built-in long-range decay: far-apart tokens attend more weakly, the inductive bias you want.1THE THREE POSITIONAL SCHEMES23 scheme how position enters relative? extend context?4 ------------- ------------------------ --------- -----------------------------5 sinusoidal add fixed sin/cos vector approx degrades past train length6 learned-abs add trainable PE[pos] no NO row past max_len -> retrain7 RoPE (rotary) rotate Q,K by m*theta YES (exact) YES via base/freq scaling8 ALiBi linear distance bias YES strong length extrapolation910 RoPE identity: = -> attention sees only (m - n)11 Cost of RoPE: 0 params, ~2 mults + 1 trig per dim-pair. Long-range decay is free.1import torch23# RoPE: rotate each (even,odd) dim-pair of q/k by an angle that scales with position.4def rope(x, base=10000.0):5 # x: (B, T, n_head, d_head), d_head even6 B, T, H, D = x.shape7 pos = torch.arange(T).float()[:, None] # (T,1)8 i = torch.arange(0, D, 2).float() # pair index9 theta = base ** (-i / D) # (D/2,)10 ang = pos * theta # (T, D/2) = m * theta_i11 cos, sin = ang.cos(), ang.sin()12 x1, x2 = x[..., 0::2], x[..., 1::2] # even, odd dims13 # 2x2 rotation per pair:14 xr1 = x1 * cos[None, :, None, :] - x2 * sin[None, :, None, :]15 xr2 = x1 * sin[None, :, None, :] + x2 * cos[None, :, None, :]16 return torch.stack([xr1, xr2], dim=-1).flatten(-2) # back to (B,T,H,D)1718# apply to q and k BEFORE the dot product; never to v. 0 learned params.Key idea
Common mistake
“Positional encoding is added to the value vectors too, so position flows through the whole layer.”
How position caps (and extends) context length
s = L_new/L_train (preserves the angle distribution but compresses high-frequency components). NTK-aware scaling rescales the 10000 base instead, preserving high frequencies and degrading low ones. YaRN combines NTK-by-parts interpolation (rescale per-dimension by wavelength) with an attention-temperature factor, and reaches 100k+ context with roughly 10× fewer training tokens and ~2.5× fewer steps than PI. ALiBi takes a different route entirely — a linear distance penalty on attention scores — and extrapolates to longer sequences strongly without any learned position.Common mistake
“A bigger context window is just a config flag — set max_len higher and you’re done.”
ALiBi, NoPE & the long-context frontier
Interview prep
- 01“Why do transformers need positional encoding?” → self-attention is permutation-invariant; without position, order is invisible.
- 02“How does RoPE work?” → rotate Q,K by m·θ per dimension-pair; the post-rotation dot product depends only on the relative offset m−n.
- 03“Why is RoPE called relative if it’s applied absolutely?” → the identity ⟨R_m q, R_n k⟩ = ⟨q, R_(n−m) k⟩ — absolute rotation, relative effect.
- 04“RoPE vs sinusoidal vs learned-absolute?” → absolute schemes cap at max_len and need retraining; RoPE has 0 params and extends via base/freq scaling.
- 05“What does the 10000 base / frequency control?” → the wavelength spread; rescaling it (NTK/YaRN) is how you extend context.
- 06“How do you extend a model to 128k context?” → PI / NTK-aware / YaRN scaling plus a little fine-tuning; validate with needle-in-a-haystack, not perplexity.
- 07“ALiBi vs RoPE for long context?” → ALiBi’s linear distance bias extrapolates strongly with no learned position; RoPE is the de-facto default and scales well with YaRN.
- 08“LLM embedding vs retrieval embedding?” → per-token (B,T,d) inside the LLM vs pooled, normalized single vector for cosine similarity in retrieval.
Common mistake
The red-flag answer: “RoPE is just a better positional encoding because the math is nicer.”
Checkpoint
You take a model trained at 4k context, set its max position to 32k with no other change, and ship. Perplexity on your eval set barely moves, but users report it “forgets” facts pasted near the top of long inputs. Best read?
Checkpoint
An interviewer asks why RoPE is described as a “relative” position scheme even though each token is rotated by its absolute index. Best answer?
Checkpoint
A retrieval engineer asks why they can’t just feed your decoder LLM’s internal token vectors into a vector DB for semantic search. Most accurate answer?
Checkpoint
Two candidate models: model A uses a learned absolute position table trained to 2k; model B uses RoPE. A product needs occasional 16k-token inputs with light fine-tuning budget. Which is the safer base and why?
Checkpoint
You shuffle the order of tokens in a prompt (keeping the same tokens) and, with positional encoding removed, the attention scores are unchanged. What property is this demonstrating?
Could you explain the RoPE identity, contrast it with absolute schemes, and design a validated context-extension plan?
Takeaways
- Self-attention is permutation-invariant, so position must be injected somewhere — it’s architectural, not cosmetic.
- Absolute schemes (sinusoidal, learned) cap at max_len and need retraining to extend; learned tables have no row past their trained length.
- RoPE rotates Q,K by position so ⟨R_m q, R_n k⟩ depends only on m−n: relative effect, zero params, built-in long-range decay.
- Extend context with PI / NTK / YaRN (YaRN: ~10× fewer tokens than PI) — and validate with needle-in-a-haystack, not perplexity.
- LLM token embeddings (per-token) ≠ retrieval embeddings (pooled, normalized) — same word, different job.
Next: the transformer block & stacking — residuals, LayerNorm vs RMSNorm, pre- vs post-norm, the FFN, and why depth works.
Sources
Free to read · better with Enzo
Learn it with Enzo
Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.