Lesson 1 of 7 · 46 min

Tokenization deep dive

The model never sees text — it sees token ids from a byte-level BPE vocabulary. How merges are learned, why 50,257 is not arbitrary, why “strawberry” and arithmetic break at the tokenizer (not the network), the 4:1 chars-to-tokens economics, and the tokenizer-level failure modes that no fine-tune can fix.

The boundary you keep blaming the model for

Before a single matmul runs, your text is shredded into token ids by a tokenizer the model had no say in. That tokenizer is a frozen, learned compression scheme — and it silently decides your cost, your effective context length, and which tasks the network can ever learn. The famous failures — “how many r’s in strawberry,” shaky digit arithmetic, code that fragments weirdly — are not reasoning failures. They are tokenizer artifacts, baked in at vocab-build time. This lesson is the ground floor: get it wrong and every later cost/latency/quality argument is built on sand.
Modern GPT-style models use byte-level byte-pair encoding (BPE). The base alphabet is not letters — it is the 256 raw UTF-8 bytes. Training starts from those 256 single-byte tokens, then repeatedly scans the corpus, finds the most frequent adjacent pair of tokens, merges it into a new symbol, and records that merge. Repeat until you hit a target vocabulary size. The final vocab is the 256 base bytes plus one new id per merge. GPT-2’s vocab is 50,257 = 256 bytes + ~50,000 learned merges + one special end-of-text token. So 50,257 is not a magic number; it is a budget of how many merges you paid for.
The decode-time algorithm is the mirror image of training: take the input bytes and apply the recorded merges strictly in learned order until no merge applies. That ordering matters — BPE is greedy and order-dependent, which is why two visually similar strings can tokenize to wildly different lengths. The crucial mental model: a tokenizer is a deterministic, learned dictionary of byte fragments shaped by whatever the training corpus happened to emit frequently — not a principled linguistic segmenter.
Let’s build the GPT TokenizerAndrej Karpathy

Why byte-level, and what it costs you

Starting from 256 bytes (not Unicode code points or characters) gives one enormous property: zero out-of-vocabulary, ever. Any string in any language, any emoji, any binary blob, is representable because the worst case is falling back to single bytes. That is why byte-level BPE beat word/character vocabularies. But it has a sharp cost: the smallest unit the model can ever name is one byte, and most non-ASCII characters are multiple UTF-8 bytes. So a Latin letter like e is often one token, while a single CJK character can be 2–3 tokens, and rare emoji more. Non-English text therefore costs more tokens per character — quietly inflating both price and the share of your context window it consumes.
The number every engineer should memorize: for ordinary English, ~4 characters ≈ 1 token (~0.75 words/token). So ~4,000 characters of prose is ~1,000–1,200 tokens. Code, JSON, and non-English text tokenize less efficiently — more tokens per character — because their frequent substrings differ from the natural-language corpus the merges were learned on. Interview angle. “Estimate the token count of this 10-page document” is a back-of-envelope screen; the strong answer states the 4:1 rule, flags that code/JSON/non-English run hotter, and notes that tokenizer choice (tiktoken vs SentencePiece) shifts the ratio.
python
1import tiktoken23enc = tiktoken.get_encoding("cl100k_base")   # GPT-4-class byte-level BPE45def toks(s): return enc.encode(s)67# numbers fragment on corpus-frequency boundaries, NOT by digit grouping:8print(toks("127"))     # often ONE token  (a common number)9print(toks("677"))     # often TWO tokens (less common -> splits)10print(toks("12834"))   # often THREE tokens1112# the spelling trap: the model never sees the letters of "strawberry"13print(toks("strawberry"))   # a couple of sub-word fragments, not 10 letters14# so "count the r's" fights the tokenizer -- give the model a TOOL, not a bigger net.1516# leading-space matters: " the" and "the" are usually DIFFERENT ids.17print(toks(" the"), toks("the"))
Look hard at the numbers above. 127 is one token but 677 is two — the boundary is set by training-corpus frequency, not by any digit-grouping rule. That inconsistency is precisely why per-digit arithmetic is fragile: the same digit 7 lives glued to different neighbors across many tokens, so the network can never learn a clean column-by-column addition algorithm. This is also why some labs (and the Llama tokenizer) deliberately split numbers into single digits — a tokenizer policy choice that measurably changes arithmetic ability, per the Hugging Face number-tokenization study.

The tokenizer→model boundary, in exact shapes

The tokenizer emits an integer tensor of shape (B, T) — batch by sequence length in tokens. The first thing the model does is a lookup into the token embedding table of shape (vocab_size, d_model), producing (B, T, d_model) (e.g. (B, T, 768) for GPT-2 124M). That (B, T, 768) tensor is what flows into the first transformer block. Everything downstream — attention, the FFN, the KV cache, the context-window limit, the per-token price — is denominated in this T. Tokens are the universal currency of the stack.
code
1WHERE TOKENIZATION SITS IN THE STACK23  "strawberry jam"4        |  tokenizer (byte-level BPE, frozen)5        v6  ids:  [302, 1618, 11, 2624]            shape (B, T)      <- integers7        |  embedding table lookup  (vocab_size x d_model)8        v9  vecs: [[...768...], [...], [...], ...]  shape (B, T, 768) <- floats10        |  + positional info, then block 1, block 2, ...11        v12  ...the rest of the transformer1314  T (token count) sets: context limit, cost, latency, KV-cache size.15  4 chars ~= 1 token (English); code/JSON/non-English run hotter.
A subtle production point interviewers love: the embedding table is tied to the output projection (the “unembedding”) in most models — the same (vocab_size, d_model) matrix is reused to map the final hidden state back to logits over the vocabulary. For GPT-2 124M that table alone is 50,257 × 768 ≈ 38.6M parameters — a third of the model; for GPT-3 175B the input embedding is 50,257 × 12,288 ≈ 617M parameters. Weight tying halves that cost and tends to improve quality. So vocabulary size is not free: every extra token is a full d_model row in two places.

Tokenizer-level failure modes that ship to prod

These are the bugs that survive every prompt tweak because they live in the vocab, not the prompt: (1) Numbers — inconsistent splits wreck arithmetic and any task that depends on exact digit manipulation. (2) Rare characters / scripts — a CJK, Arabic, or rare-emoji-heavy input can cost 2–4× the tokens of equivalent English, blowing context budgets and cost on multilingual workloads. (3) Code — identifiers, whitespace runs, and unusual symbols fragment, so a code model spends context inefficiently unless its tokenizer was trained on code. (4) “Glitch tokens” — under-trained tokens (e.g. the infamous SolidGoldMagikarp Reddit-username token) that appear in the vocab but almost never in training data, producing bizarre, unsteerable outputs when summoned.
The senior intervention when accuracy is domain-bound: inspect the token ids of representative inputs (with tokenizer.encode) before assuming character-level behavior holds, and consider retraining BPE on your target domain. Stanford’s CS336 has students implement BPE from scratch precisely so they can reason about these failure modes. Interview angle. “Why is your model trained on English Wikipedia bad at multilingual code-switching?” → non-aligned token boundaries: the merges optimize English, so other scripts fragment, costing context and degrading perplexity. A domain-tuned tokenizer is often the highest-ROI fix — higher than swapping the model.
code
1TOKENIZER FAILURE MODES (all baked in at vocab-build time)23  Symptom                         Root cause                       Real fix4  -----------------------------   ------------------------------   ---------------------------5  "count the r's" wrong           no single-token char handle      tool (code exec)6  digit arithmetic shaky          numbers split by frequency       digit-split tokenizer / tool7  multilingual cost 2-4x          non-ASCII = multi-byte tokens    domain/multilingual BPE8  code context wasted             identifiers/whitespace split     code-trained tokenizer9  bizarre output on rare string   "glitch" under-trained token     filter / retrain vocab1011  None of these are fixed by a bigger model or a better prompt.

Token economics: the 4:1 rule, in production

Tokens are the unit of everything commercial: pricing, latency, rate limits, and the context-window cap are all counted in tokens, so token accounting is the first thing a production client does. The pricing fact that flips junior intuition: output tokens cost ~3–5× input tokens on most providers, because decode is sequential and compute-heavy while prefill is parallelized (Lesson 5 makes this mechanical). So the cheapest feature is a big prompt with a short output — exactly backwards from “shorter prompts are cheaper.” And because the tokenizer sets the chars→tokens ratio, an inefficient tokenizer on your domain quietly inflates every one of those line items at once.
Concrete sizing you can defend in an interview. A 10-page English document (~20,000 characters) is ~5,000 tokens at the 4:1 rule. The same content as dense JSON or source code can be 1.5–2× that, because braces, quotes, indentation runs, and identifiers fragment. Feed 200,000 such requests/day through a mid-tier model at ~$2.50/M input and ~$10/M output and you are reasoning in the thousands-of-dollars-per-day range before any caching — which is why teams obsess over token count per request, not just request count. The tokenizer is the term that multiplies through all of it.
The tokenizer is the cheapest thing to get wrong and the most expensive thing to leave wrong: it sets your token bill, your effective context length, and the ceiling on every character- and digit-level task — all before the first matmul.

Interview prep

Tokenization has quietly become a coding round, not a one-line theory question — Perplexity and others ask candidates to implement a byte-level BPE tokenizer with unit tests and an O(n) streaming mode. Verbally, lead with the mechanism (merges on 256 bytes), then tie every quirk to a real failure mode and a number.
  1. 01“Explain BPE.” → start from 256 UTF-8 bytes, iteratively merge the most frequent adjacent pair, record merges; decode replays them in order. Vocab = 256 + #merges.
  2. 02“Why is GPT-2’s vocab 50,257?” → 256 base bytes + ~50,000 merges + 1 end-of-text token; the merge count is a budget you choose.
  3. 03“Why are LLMs bad at counting letters / arithmetic?” → no single-token char handle; numbers fragment by corpus frequency. Give it a tool or split digits.
  4. 04“Estimate tokens for this doc.” → ~4 chars/token for English; code, JSON, and non-English run hotter. State the assumption.
  5. 05“Byte-level BPE (GPT-2) vs SentencePiece (Llama)?” → both subword; byte-level = zero OOV on raw bytes; SentencePiece works on raw text with a unigram/BPE model and explicit space handling.
  6. 06“Why does an English-trained model struggle on multilingual code?” → non-aligned token boundaries; other scripts fragment, costing context and perplexity.
  7. 07“How would you choose vocab size?” → trade compression (fewer tokens/seq → cheaper, longer effective context) vs a bigger, costlier embedding/unembedding table and rarer-token under-training.
  8. 08“What’s a glitch token?” → an under-trained vocab entry (e.g. SolidGoldMagikarp) that triggers unsteerable output; filter it or retrain the tokenizer.
Going deeper, the follow-ups that separate seniors: “implement BPE from scratch” (build merge rules, handle special tokens, add an O(n) streaming encode — naive dict-of-frequencies fails the time/memory budget); “why does leading whitespace change tokenization?” (" the" and "the" are different ids, which is why prompt formatting subtly shifts behavior); “what does tiktoken do that SentencePiece doesn’t?” (tiktoken is byte-level BPE with a fast Rust core and a regex pre-tokenizer; SentencePiece is a self-contained trainer/encoder with unigram-LM and reversible space markers). Always cite the number-tokenization evidence that how numbers are tokenized is a real, measurable contributor to arithmetic ability.
repominbpe — minimal, clean byte-pair-encoding codeAndrej KarpathypaperNeural Machine Translation of Rare Words with Subword Units (the BPE paper)Sennrich et al. (arXiv)articleNumber Tokenization — how digit-splitting changes arithmeticHugging FacedocsTiktokenizer — see exactly how text becomes tokensinteractive playground

Checkpoint

Your model nails “count the letters in ELEPHANT” when you let it write Python, but fails when asked directly. A teammate wants to fine-tune on letter-counting examples. Best read?

AFine-tune on many letter-counting examples — the model just hasn’t seen enoughBIt’s a tokenization artifact: the model never sees characters, so route character/number tasks to a tool (code execution) instead of fine-tuningCSwitch to a larger model with a bigger context window
Sign up free to answer and see why

Checkpoint

A multilingual support feature is 3× over budget on cost and frequently truncates context — but only for Japanese and Arabic tickets. English is fine. Most likely cause?

AThe model is worse at non-English languages, so it retries moreBA rate limit is throttling those regionsCByte-level BPE trained mostly on English fragments non-ASCII scripts into 2–4× more tokens, inflating cost and consuming the context budget
Sign up free to answer and see why

Checkpoint

You build a code-assistant on a model whose tokenizer was trained on web prose. Retrieval and prompting are tuned, yet long files blow the context window sooner than expected. Strongest lever?

AAdopt (or fine-tune) a tokenizer trained on code so identifiers and whitespace runs don’t fragment into many tokensBIncrease temperature so the model is more conciseCSummarize each file before sending it
Sign up free to answer and see why

Checkpoint

An interviewer asks you to size the embedding table for a 50k-vocab, d_model=4096 model and comment on cost. Best answer?

AIt’s negligible — embeddings are tiny compared to attentionBAbout 50,000 × 4,096 ≈ 205M parameters, and it’s typically tied to the output unembedding so you don’t pay for it twiceCYou can’t estimate it without knowing the number of layers
Sign up free to answer and see why

Checkpoint

A model emits bizarre, unsteerable text whenever a specific rare username-like string appears in the prompt — even across temperatures and rephrasings. Best explanation?

APrompt injection from the userBA sampling bug at low temperatureCA “glitch token”: an under-trained vocab entry (like SolidGoldMagikarp) that the model has almost no learned representation for, so its output is unstable
Sign up free to answer and see why

Could you implement byte-level BPE, estimate token counts for any input, and explain every tokenizer-level failure mode with its real fix?

New to itGetting thereConfident

Takeaways

  • BPE merges the most frequent adjacent pair starting from 256 UTF-8 bytes; vocab = 256 + #merges (GPT-2: 50,257).
  • ~4 chars ≈ 1 token (English); code, JSON, and non-English run 2–4× hotter — that’s your cost and context budget.
  • Spelling and arithmetic failures are tokenizer artifacts, not reasoning failures — use a tool or a digit-splitting tokenizer.
  • The embedding table is vocab × d_model (often tied to the unembedding); vocabulary size is a real parameter cost.
  • Tokenizer-level failures (numbers, scripts, code, glitch tokens) survive every prompt tweak — fix the vocab, not the prompt.

Next: attention from scratch — scaled dot-product, why we divide by √d_k, multi-head, exact shapes, and the quadratic cost.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.