Tokenization splits raw text into integer IDs; it determines vocabulary size, sequence length, OOV handling, and downstream cost (more tokens = more money and latency). BPE (Byte Pair Encoding) iteratively merges the most frequent adjacent pairs, starting from bytes; it's what GPT-2/3/4 use and handles any UTF-8 cleanly. WordPiece (BERT-family) is similar but uses likelihood-based merges. SentencePiece is a system, not an algorithm: it can run BPE or Unigram language-model tokenization directly on raw text without pre-tokenization, which makes it ideal for languages without whitespace (CJK, Thai). Unigram (also via SentencePiece) starts with a large vocabulary and prunes; it allows multiple segmentations and probabilistic decoding, useful for languages with rich morphology. For a modern multilingual LLM I'd pick SentencePiece + BPE for simplicity, or Unigram if I cared about morphological precision and had the compute to tune it. Senior points: tokenization quality affects arithmetic ("3.11 vs 3.9"), code ("\n" handling), and multilingual fairness; modern models are investigating token-free approaches like ByT5 or Mamba-Bytes for low-resource languages.