Lesson 1 of 7 · 46 min
Tokenization deep dive
The model never sees text — it sees token ids from a byte-level BPE vocabulary. How merges are learned, why 50,257 is not arbitrary, why “strawberry” and arithmetic break at the tokenizer (not the network), the 4:1 chars-to-tokens economics, and the tokenizer-level failure modes that no fine-tune can fix.
The boundary you keep blaming the model for
Let’s build the GPT TokenizerAndrej KarpathyWhy byte-level, and what it costs you
e is often one token, while a single CJK character can be 2–3 tokens, and rare emoji more. Non-English text therefore costs more tokens per character — quietly inflating both price and the share of your context window it consumes.1import tiktoken23enc = tiktoken.get_encoding("cl100k_base") # GPT-4-class byte-level BPE45def toks(s): return enc.encode(s)67# numbers fragment on corpus-frequency boundaries, NOT by digit grouping:8print(toks("127")) # often ONE token (a common number)9print(toks("677")) # often TWO tokens (less common -> splits)10print(toks("12834")) # often THREE tokens1112# the spelling trap: the model never sees the letters of "strawberry"13print(toks("strawberry")) # a couple of sub-word fragments, not 10 letters14# so "count the r's" fights the tokenizer -- give the model a TOOL, not a bigger net.1516# leading-space matters: " the" and "the" are usually DIFFERENT ids.17print(toks(" the"), toks("the"))127 is one token but 677 is two — the boundary is set by training-corpus frequency, not by any digit-grouping rule. That inconsistency is precisely why per-digit arithmetic is fragile: the same digit 7 lives glued to different neighbors across many tokens, so the network can never learn a clean column-by-column addition algorithm. This is also why some labs (and the Llama tokenizer) deliberately split numbers into single digits — a tokenizer policy choice that measurably changes arithmetic ability, per the Hugging Face number-tokenization study.Key idea
The tokenizer→model boundary, in exact shapes
(B, T) — batch by sequence length in tokens. The first thing the model does is a lookup into the token embedding table of shape (vocab_size, d_model), producing (B, T, d_model) (e.g. (B, T, 768) for GPT-2 124M). That (B, T, 768) tensor is what flows into the first transformer block. Everything downstream — attention, the FFN, the KV cache, the context-window limit, the per-token price — is denominated in this T. Tokens are the universal currency of the stack.1WHERE TOKENIZATION SITS IN THE STACK23 "strawberry jam"4 | tokenizer (byte-level BPE, frozen)5 v6 ids: [302, 1618, 11, 2624] shape (B, T) <- integers7 | embedding table lookup (vocab_size x d_model)8 v9 vecs: [[...768...], [...], [...], ...] shape (B, T, 768) <- floats10 | + positional info, then block 1, block 2, ...11 v12 ...the rest of the transformer1314 T (token count) sets: context limit, cost, latency, KV-cache size.15 4 chars ~= 1 token (English); code/JSON/non-English run hotter.(vocab_size, d_model) matrix is reused to map the final hidden state back to logits over the vocabulary. For GPT-2 124M that table alone is 50,257 × 768 ≈ 38.6M parameters — a third of the model; for GPT-3 175B the input embedding is 50,257 × 12,288 ≈ 617M parameters. Weight tying halves that cost and tends to improve quality. So vocabulary size is not free: every extra token is a full d_model row in two places.Tokenizer-level failure modes that ship to prod
SolidGoldMagikarp Reddit-username token) that appear in the vocab but almost never in training data, producing bizarre, unsteerable outputs when summoned.tokenizer.encode) before assuming character-level behavior holds, and consider retraining BPE on your target domain. Stanford’s CS336 has students implement BPE from scratch precisely so they can reason about these failure modes. Interview angle. “Why is your model trained on English Wikipedia bad at multilingual code-switching?” → non-aligned token boundaries: the merges optimize English, so other scripts fragment, costing context and degrading perplexity. A domain-tuned tokenizer is often the highest-ROI fix — higher than swapping the model.1TOKENIZER FAILURE MODES (all baked in at vocab-build time)23 Symptom Root cause Real fix4 ----------------------------- ------------------------------ ---------------------------5 "count the r's" wrong no single-token char handle tool (code exec)6 digit arithmetic shaky numbers split by frequency digit-split tokenizer / tool7 multilingual cost 2-4x non-ASCII = multi-byte tokens domain/multilingual BPE8 code context wasted identifiers/whitespace split code-trained tokenizer9 bizarre output on rare string "glitch" under-trained token filter / retrain vocab1011 None of these are fixed by a bigger model or a better prompt.Common mistake
“A token is basically a word, so character and number tasks should be easy.”
Token economics: the 4:1 rule, in production
The tokenizer is the cheapest thing to get wrong and the most expensive thing to leave wrong: it sets your token bill, your effective context length, and the ceiling on every character- and digit-level task — all before the first matmul.
Interview prep
- 01“Explain BPE.” → start from 256 UTF-8 bytes, iteratively merge the most frequent adjacent pair, record merges; decode replays them in order. Vocab = 256 + #merges.
- 02“Why is GPT-2’s vocab 50,257?” → 256 base bytes + ~50,000 merges + 1 end-of-text token; the merge count is a budget you choose.
- 03“Why are LLMs bad at counting letters / arithmetic?” → no single-token char handle; numbers fragment by corpus frequency. Give it a tool or split digits.
- 04“Estimate tokens for this doc.” → ~4 chars/token for English; code, JSON, and non-English run hotter. State the assumption.
- 05“Byte-level BPE (GPT-2) vs SentencePiece (Llama)?” → both subword; byte-level = zero OOV on raw bytes; SentencePiece works on raw text with a unigram/BPE model and explicit space handling.
- 06“Why does an English-trained model struggle on multilingual code?” → non-aligned token boundaries; other scripts fragment, costing context and perplexity.
- 07“How would you choose vocab size?” → trade compression (fewer tokens/seq → cheaper, longer effective context) vs a bigger, costlier embedding/unembedding table and rarer-token under-training.
- 08“What’s a glitch token?” → an under-trained vocab entry (e.g. SolidGoldMagikarp) that triggers unsteerable output; filter it or retrain the tokenizer.
" the" and "the" are different ids, which is why prompt formatting subtly shifts behavior); “what does tiktoken do that SentencePiece doesn’t?” (tiktoken is byte-level BPE with a fast Rust core and a regex pre-tokenizer; SentencePiece is a self-contained trainer/encoder with unigram-LM and reversible space markers). Always cite the number-tokenization evidence that how numbers are tokenized is a real, measurable contributor to arithmetic ability.Common mistake
The red-flag answer: “tokenization is just a preprocessing detail — the model figures out the rest.”
Checkpoint
Your model nails “count the letters in ELEPHANT” when you let it write Python, but fails when asked directly. A teammate wants to fine-tune on letter-counting examples. Best read?
Checkpoint
A multilingual support feature is 3× over budget on cost and frequently truncates context — but only for Japanese and Arabic tickets. English is fine. Most likely cause?
Checkpoint
You build a code-assistant on a model whose tokenizer was trained on web prose. Retrieval and prompting are tuned, yet long files blow the context window sooner than expected. Strongest lever?
Checkpoint
An interviewer asks you to size the embedding table for a 50k-vocab, d_model=4096 model and comment on cost. Best answer?
Checkpoint
A model emits bizarre, unsteerable text whenever a specific rare username-like string appears in the prompt — even across temperatures and rephrasings. Best explanation?
Could you implement byte-level BPE, estimate token counts for any input, and explain every tokenizer-level failure mode with its real fix?
Takeaways
- BPE merges the most frequent adjacent pair starting from 256 UTF-8 bytes; vocab = 256 + #merges (GPT-2: 50,257).
- ~4 chars ≈ 1 token (English); code, JSON, and non-English run 2–4× hotter — that’s your cost and context budget.
- Spelling and arithmetic failures are tokenizer artifacts, not reasoning failures — use a tool or a digit-splitting tokenizer.
- The embedding table is vocab × d_model (often tied to the unembedding); vocabulary size is a real parameter cost.
- Tokenizer-level failures (numbers, scripts, code, glitch tokens) survive every prompt tweak — fix the vocab, not the prompt.
Next: attention from scratch — scaled dot-product, why we divide by √d_k, multi-head, exact shapes, and the quadratic cost.
Sources
Free to read · better with Enzo
Learn it with Enzo
Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.