Lesson 3 of 7 · 47 min

Embeddings & positional encoding

What the embedding vector actually encodes, how encoders pool tokens into one vector, and the position problem: attention is permutation-invariant, so order must be injected. Sinusoidal vs learned-absolute vs RoPE — the rotation identity that makes RoPE relative — and how the positional scheme caps your context length.

Attention can’t tell “dog bites man” from “man bites dog”

Self-attention is permutation-invariant: shuffle the input tokens and the set of attention scores is identical, because the dot product q·k doesn’t know where either token sits. That’s a problem — language is ordered. So every transformer injects position somewhere, and the scheme you pick is not cosmetic: it decides whether the model can run at a longer context than it trained on, with zero extra parameters or a brittle retrain. This lesson covers what embeddings encode and the three positional schemes you’ll be grilled on, ending with why RoPE won.
Start with the token embedding. It’s a lookup into a (vocab_size, d_model) table that maps each token id to a learned d_model-vector. There’s nothing magic in a single row at init — meaning emerges from training: tokens that play similar roles get pulled to nearby points, so the geometry of the embedding space comes to encode syntactic and semantic relationships. This is the same object as a word2vec/GloVe embedding, except it’s learned jointly with the rest of the network and is contextualized by later layers (the embedding of a token is the same on the way in; what attention produces downstream is context-dependent).
A point that trips people up: a decoder LLM’s input embedding is not the same thing as a sentence embedding from a retrieval model. Retrieval encoders pool the per-token vectors into one fixed-length vector — mean-pooling or a special CLS token — so a 3-word query and a 400-word passage land in the same space and are comparable by cosine similarity (the embeddings track from the RAG track lives here). A generative LLM keeps a per-token (B, T, d_model) tensor all the way through and never pools. Interview angle. “What’s the difference between the embedding inside GPT and an embedding from text-embedding-3?” → per-token vs pooled-and-normalized for similarity; same word “embedding,” different job.
What does the geometry actually encode? Two things worth being precise about. Direction carries meaning — the classic king − man + woman ≈ queen demonstrates that semantic relations show up as consistent vector offsets, which is why cosine similarity (angle, not magnitude) is the right comparison. Magnitude is mostly a frequency/usage artifact, which is exactly why retrieval normalizes vectors before comparing. Inside an LLM the input embedding is the same vector for a token regardless of context; the contextual meaning (“bank” the riverside vs the financial institution) is built up by the attention layers that follow, not by the embedding row itself. That separation — static lookup, then contextualization — is the mental model to carry.
One more production detail: the embedding dimension d_model is a capacity knob with a real cost. GPT-2 124M uses d_model=768; GPT-3 175B uses 12,288. Because the embedding table is vocab × d_model and is typically tied to the unembedding (Lesson 1), widening d_model grows two of the largest matrices in the model and every activation tensor downstream. So “just use bigger embeddings” is not free — it scales the whole network’s width. Interview angle. “Why not make d_model huge for richer representations?” → it widens every layer’s matmuls and activations and the tied embedding/unembedding tables; depth and attention design usually buy more than raw width.
RoPE: Understanding Rotary Positional Embeddings in transformersHugging Face

Sinusoidal absolute encodings (the original)

Vaswani et al. injected position with a fixed table of sinusoids added elementwise to the token embedding: PE(pos, 2i) = sin(pos / 10000^(2i/d_model)) and PE(pos, 2i+1) = cos(pos / 10000^(2i/d_model)). Each position 0,1,2,… gets a unique d_model-vector, with frequencies in a geometric progression from short to very long wavelengths. The reason for sinusoids over a plain integer index: any shifted position PE(pos+k) is a linear function of PE(pos), which is the structure a model needs to learn relative offsets. d2l.ai notes that learned absolute embeddings (a trainable (max_len, d_model) table, as GPT-2 uses) perform comparably — but both share a hard limit.
That hard limit is the crux: absolute schemes (sinusoidal or learned) bind position into the embedding up to some max_len. A learned table has literally no row for position 4097 if you trained to 4096; sinusoidal can be evaluated past training length but the model never saw those phase combinations and degrades. So with absolute encodings, extending context means retraining — and that’s the pain RoPE was designed to remove.

RoPE — rotate Q and K by position

Rotary Position Embedding (RoFormer, Su et al. 2021) is the scheme nearly every 2023–2026 open LLM uses (Llama, Mistral, Gemma, Qwen). Instead of adding a position vector, it rotates each query and key vector by an angle that depends on its position. Group the d dimensions into d/2 pairs; pair i is rotated by angle m·θᵢ at position m, where θᵢ = 10000^(−2(i−1)/d) (the same 10000 base as sinusoids). Concretely it’s a block-diagonal matrix of 2×2 rotations applied to q and k before the dot product — no learned parameters, two multiplies and a sin/cos per pair, near-zero FLOPs.
Here’s the identity that makes it brilliant. After rotating, the inner product of a query at position m and a key at position n depends only on the relative offset (m − n): ⟨R_m·q, R_n·k⟩ = ⟨q, R_(n−m)·k⟩. So RoPE encodes position absolutely (each token is rotated by its own index) yet makes attention behave relatively (only distance matters) — with zero extra parameters and zero learned position table. And because high-frequency pairs decorrelate over long distances, RoPE has a built-in long-range decay: far-apart tokens attend more weakly, the inductive bias you want.
code
1THE THREE POSITIONAL SCHEMES23  scheme          how position enters        relative?   extend context?4  -------------   ------------------------   ---------   -----------------------------5  sinusoidal      add fixed sin/cos vector   approx      degrades past train length6  learned-abs     add trainable PE[pos]      no          NO row past max_len -> retrain7  RoPE (rotary)   rotate Q,K by m*theta      YES (exact) YES via base/freq scaling8  ALiBi           linear distance bias       YES         strong length extrapolation910  RoPE identity:   =    -> attention sees only (m - n)11  Cost of RoPE:   0 params, ~2 mults + 1 trig per dim-pair. Long-range decay is free.
python
1import torch23# RoPE: rotate each (even,odd) dim-pair of q/k by an angle that scales with position.4def rope(x, base=10000.0):5    # x: (B, T, n_head, d_head), d_head even6    B, T, H, D = x.shape7    pos = torch.arange(T).float()[:, None]                      # (T,1)8    i = torch.arange(0, D, 2).float()                           # pair index9    theta = base ** (-i / D)                                    # (D/2,)10    ang = pos * theta                                           # (T, D/2) = m * theta_i11    cos, sin = ang.cos(), ang.sin()12    x1, x2 = x[..., 0::2], x[..., 1::2]                         # even, odd dims13    # 2x2 rotation per pair:14    xr1 = x1 * cos[None, :, None, :] - x2 * sin[None, :, None, :]15    xr2 = x1 * sin[None, :, None, :] + x2 * cos[None, :, None, :]16    return torch.stack([xr1, xr2], dim=-1).flatten(-2)         # back to (B,T,H,D)1718# apply to q and k BEFORE the dot product; never to v. 0 learned params.

How position caps (and extends) context length

Because RoPE’s frequencies are tied to position, you can extend context after training by adjusting them — the family of techniques every long-context model uses. Position Interpolation (PI) linearly rescales position indices by s = L_new/L_train (preserves the angle distribution but compresses high-frequency components). NTK-aware scaling rescales the 10000 base instead, preserving high frequencies and degrading low ones. YaRN combines NTK-by-parts interpolation (rescale per-dimension by wavelength) with an attention-temperature factor, and reaches 100k+ context with roughly 10× fewer training tokens and ~2.5× fewer steps than PI. ALiBi takes a different route entirely — a linear distance penalty on attention scores — and extrapolates to longer sequences strongly without any learned position.
The failure mode that bites teams: naive context extension degrades retrieval before it degrades perplexity. A model with no-finetune RoPE scaling can still post a healthy next-token loss while silently losing the ability to find a fact buried at position 90k — and standard perplexity evals miss it. Interview angle. “You stretched a 4k model to 32k with RoPE scaling — how do you know it works?” The strong answer: rerun a needle-in-a-haystack probe (plant a fact at varied depths, measure recall), not just perplexity; prefer YaRN over NTK over PI; and retrain on a small slice of long-context data rather than shipping a drop-in. Saying “perplexity looked fine” is the tell of someone who hasn’t shipped long context.

ALiBi, NoPE & the long-context frontier

RoPE is the default, but two alternatives come up. ALiBi (Attention with Linear Biases) skips position vectors entirely and instead subtracts a distance-proportional penalty from the attention scores — closer tokens get a smaller penalty, farther ones a larger one, with a fixed per-head slope. It has zero learned position parameters and extrapolates to far longer sequences than it trained on remarkably well, which is why some long-context models prefer it. NoPE (No Positional Encoding) is the surprising result that a sufficiently trained decoder can infer position from the causal mask alone in some regimes — evidence that order information leaks through the autoregressive structure even without an explicit scheme.
The frontier is moving fast and the honest senior answer names the open questions. Combining RoPE scaling (YaRN) with Mistral-style sliding-window attention at 1M-token contexts is live research, not settled practice. LongRoPE2 and similar work push toward near-lossless scaling. And crucially, none of these position tricks touch the KV-cache cost — a 1M-token context is a memory problem (Lessons 5–6) regardless of how you encode position. Interview angle. “How would you get a model to 1M tokens?” → it is not one lever: position scaling (YaRN) and attention sparsity (sliding window) and KV-cache management (paging/quantization) and a needle-in-a-haystack gate — name all four, because any one alone fails.

Interview prep

Positional encoding is where the most strong-vs-weak splits happen in internals rounds. The weak answer is “RoPE rotates queries and keys”; the strong answer explains why rotation yields relative position and ties it to long-context extrapolation (YaRN/PI). Lead with the permutation-invariance problem, then the scheme, then the context-length consequence.
  1. 01“Why do transformers need positional encoding?” → self-attention is permutation-invariant; without position, order is invisible.
  2. 02“How does RoPE work?” → rotate Q,K by m·θ per dimension-pair; the post-rotation dot product depends only on the relative offset m−n.
  3. 03“Why is RoPE called relative if it’s applied absolutely?” → the identity ⟨R_m q, R_n k⟩ = ⟨q, R_(n−m) k⟩ — absolute rotation, relative effect.
  4. 04“RoPE vs sinusoidal vs learned-absolute?” → absolute schemes cap at max_len and need retraining; RoPE has 0 params and extends via base/freq scaling.
  5. 05“What does the 10000 base / frequency control?” → the wavelength spread; rescaling it (NTK/YaRN) is how you extend context.
  6. 06“How do you extend a model to 128k context?” → PI / NTK-aware / YaRN scaling plus a little fine-tuning; validate with needle-in-a-haystack, not perplexity.
  7. 07“ALiBi vs RoPE for long context?” → ALiBi’s linear distance bias extrapolates strongly with no learned position; RoPE is the de-facto default and scales well with YaRN.
  8. 08“LLM embedding vs retrieval embedding?” → per-token (B,T,d) inside the LLM vs pooled, normalized single vector for cosine similarity in retrieval.
Going deeper: “why does scaling the base frequency help?” (it stretches the wavelengths so positions seen at train time now cover a longer range, reducing aliasing at unseen positions); “what breaks first when you over-extend?” (structured retrieval — the needle test — well before next-token loss); “can you mix RoPE with sliding-window attention at 1M context?” (an open research question; YaRN + Mistral-style windows is live work, and KV-cache cost still bites regardless of the position scheme). Note that position interacts with the tokenizer decision from Lesson 1: longer effective sequences from an inefficient tokenizer reach the context cap sooner.
paperRoFormer: Enhanced Transformer with Rotary Position EmbeddingSu et al. (arXiv)articleRotary Embeddings: A Relative RevolutionEleutherAIpaperYaRN: Efficient Context Window Extension of Large Language ModelsPeng et al. (arXiv)docsd2l.ai — Self-Attention and Positional EncodingDive into Deep Learning

Checkpoint

You take a model trained at 4k context, set its max position to 32k with no other change, and ship. Perplexity on your eval set barely moves, but users report it “forgets” facts pasted near the top of long inputs. Best read?

APerplexity is fine, so the extension worked — it’s a prompt-formatting issueBNaive context extension degraded long-range retrieval; validate with needle-in-a-haystack and use YaRN/PI scaling plus a little fine-tuningCThe context window can’t be changed after training under any scheme
Sign up free to answer and see why

Checkpoint

An interviewer asks why RoPE is described as a “relative” position scheme even though each token is rotated by its absolute index. Best answer?

ABecause it’s added to the embedding like sinusoidal encodingsBBecause the post-rotation dot product ⟨R_m q, R_n k⟩ equals ⟨q, R_(n−m) k⟩ — it depends only on the offset m−nCBecause it has learnable parameters that capture relative distance
Sign up free to answer and see why

Checkpoint

A retrieval engineer asks why they can’t just feed your decoder LLM’s internal token vectors into a vector DB for semantic search. Most accurate answer?

AThe LLM keeps a per-token (B,T,d) tensor and never pools; retrieval needs one pooled, normalized vector per text for cosine comparisonBLLM embeddings are too high-dimensional for any vector DBCDecoder models don’t have embeddings at all
Sign up free to answer and see why

Checkpoint

Two candidate models: model A uses a learned absolute position table trained to 2k; model B uses RoPE. A product needs occasional 16k-token inputs with light fine-tuning budget. Which is the safer base and why?

AModel A — learned tables generalize to any length once trainedBEither works equally; position scheme doesn’t affect context extensionCModel B — RoPE extends via base/frequency scaling (PI/NTK/YaRN) plus light fine-tuning, while the learned-absolute table has no representation past 2k
Sign up free to answer and see why

Checkpoint

You shuffle the order of tokens in a prompt (keeping the same tokens) and, with positional encoding removed, the attention scores are unchanged. What property is this demonstrating?

ANumerical instability in the softmaxBSelf-attention is permutation-invariant, which is exactly why position must be injectedCThe model has collapsed all tokens to the same embedding
Sign up free to answer and see why

Could you explain the RoPE identity, contrast it with absolute schemes, and design a validated context-extension plan?

New to itGetting thereConfident

Takeaways

  • Self-attention is permutation-invariant, so position must be injected somewhere — it’s architectural, not cosmetic.
  • Absolute schemes (sinusoidal, learned) cap at max_len and need retraining to extend; learned tables have no row past their trained length.
  • RoPE rotates Q,K by position so ⟨R_m q, R_n k⟩ depends only on m−n: relative effect, zero params, built-in long-range decay.
  • Extend context with PI / NTK / YaRN (YaRN: ~10× fewer tokens than PI) — and validate with needle-in-a-haystack, not perplexity.
  • LLM token embeddings (per-token) ≠ retrieval embeddings (pooled, normalized) — same word, different job.

Next: the transformer block & stacking — residuals, LayerNorm vs RMSNorm, pre- vs post-norm, the FFN, and why depth works.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.