The unglamorous decision that silently caps every RAG system — chunking failure modes, sizing by metric, parsing as the real bottleneck, late chunking, Anthropic’s contextual retrieval, small-to-big, and how systems like Cursor chunk at scale.
The failure nobody instruments
Most “the model hallucinated” bugs are really “the answer-bearing sentence was never retrieved in one piece.” In Barnett et al.’s study of seven production RAG failure points, wrong chunking strategy and not-in-context (the right document is found, but the relevant passage is split or buried) rank among the most common — and both are decided long before a query runs, at index time.
Chunking is the highest-leverage, least-glamorous decision in RAG. It happens once, offline, and then silently caps everything downstream: a chunk that doesn’t exist can’t be retrieved, reranked, or grounded. The core tension is simple and brutal — your unit of retrieval (what you embed) is almost never your unit of meaning (what actually answers the question). Embed too coarse and the signal washes out; embed too fine and the chunk loses the context that grounds it. Everything in this lesson is a way to escape that vice.
Interview angle. Chunking is a favourite design-round topic because it has no single right answer — it’s pure tradeoff reasoning, and that’s exactly what interviewers want to see. A strong candidate never says “I’d use 512 tokens”; they say “I’d start at ~512 with structure-aware splitting, then measure precision@5 on a labelled set and adjust, and I’d decouple the unit I match from the unit I read.” Naming the metric and the tension is the signal.
Parsing comes first — garbage chunks start as garbage parses
Before you split anything, you have to extract clean text — and for real corpora (PDFs, scanned contracts, HTML, slide decks) this is where most quality is silently lost. A naive PDF-to-text dump linearizes a two-column page into interleaved nonsense, turns a table into a wall of numbers with no row/column structure, drops headers that gave a section its meaning, and mangles equations. No chunker can recover a sentence the parser already shredded. The production rule: treat parsing as a first-class stage with its own quality bar — use layout-aware extraction (e.g. table-and-heading-preserving parsers), keep tables as Markdown or HTML so structure survives embedding, and OCR scanned pages rather than embedding the image-derived garbage.
This connects directly to the eval lesson: one documented team saw end-to-end accuracy drop after a chunking change, and the split metrics revealed recall@5 had actually risen while precision collapsed — a parser was fragmenting tables into noise. The chunker got blamed; the parser was the culprit. Tables, code, and figures are the usual victims — handle them with structure-aware splitting (below), never a token counter.
The five families — and how each one fails
01Fixed-size — split every N tokens. Trivial to build; severs sentences, tables, and code mid-thought.
02Recursive character — split on paragraph → sentence → word until it fits. The sane default for prose.
03Document-aware — split on real structure: Markdown headings, code blocks, legal clauses. Best when structure is clean.
04Semantic — start a new chunk where consecutive sentences diverge in embedding space. Powerful, but happily splits a question from its answer when the topic jumps.
05Late chunking — embed the whole document first, then pool into chunks, so each chunk vector carries document context (more below).
The practitioner consensus: start with recursive splitting on Markdown/sentence boundaries, switch to document-aware the moment your corpus has real structure (code, contracts, API docs). Reach for semantic chunking only after you’ve measured that boundary placement is your problem — it adds cost and a new failure mode (over-splitting on a topic jump) that often isn’t worth it.
Default to ~512 tokens with 10–15% overlap, recursively split on structure — then measure. The rule that survives contact with real corpora: if precision@5 < 0.7, drop toward 256 tokens. Practitioner benchmarks repeatedly find 256-token chunks beat 384 on precision, because too large dilutes the embedding (the one relevant sentence is averaged across hundreds of irrelevant tokens, similarity drops, the chunk ranks lower) while too small strips the entities and pronouns that ground a chunk — the “meaningless out of context” failure we fix below.
The mechanism behind “too large dilutes” is worth internalizing because it’s counterintuitive. A chunk is collapsed into one vector by pooling its token embeddings. If the chunk is 1,500 tokens and only 30 of them are about your query, the query-relevant signal is averaged against 1,470 tokens of unrelated text — cosine similarity to the query falls, and a tighter 200-token competitor that is all about the query outranks it. So bigger chunks don’t “add context for the model,” they add noise to the vector and push the right chunk down the ranking. That’s why the fix for “the model needs more context” is small-to-big (match small, read big), not bigger chunks.
code
1Chunk size as a precision/recall tradeoff (typical prose corpus)23 Chunk size Recall@20 Precision@5 What breaks4 ---------- --------- ----------- -----------------------------5 128 tok lower high context stripped (orphaned refs)6 256 tok good highest sweet spot for entity lookups7 512 tok high good sane default; measure before trusting8 1024 tok highest lower signal dilution; noisy top-k9 2048+ tok high poor one fact lost in an averaged vector1011 Overlap 10-15% prevents a sentence split across a boundary from vanishing.12 These shift by corpus -- the numbers are a starting point, the METRIC is the rule.
Two senior nuances on overlap and tables. Overlap (carrying ~10–15% of the previous chunk into the next) exists so a sentence straddling a boundary still lives intact in at least one chunk; too much overlap inflates your index size and re-ranks near-duplicates, so it’s a small constant, not a big one. And never token-count through a table or a code block — splitting a 5-column table at token 512 leaves two half-tables that mean nothing; keep structured units whole even when they exceed your target size, then handle the oversize ones specially.
python
1from langchain_text_splitters import RecursiveCharacterTextSplitter23splitter = RecursiveCharacterTextSplitter(4 chunk_size=512, chunk_overlap=64, # ~12% overlap5 separators=["\n## ", "\n\n", "\n", ". ", " "], # structure first6)7chunks = splitter.split_text(doc)8records = [{"text": c, "doc_id": doc_id, "chunk": i} for i, c in enumerate(chunks)]910# Don't ship a chunk size on vibes -- measure it against a labelled set.11def precision_at_k(eval_set, k=5):12 total = 0.013 for query, gold_chunk_ids in eval_set:14 got = [r["chunk_id"] for r in retrieve(query, k=k)]15 total += len(set(got) & set(gold_chunk_ids)) / k16 return total / len(eval_set)17# precision_at_k < 0.7 -> shrink chunks (and/or add a reranker), then re-measure.
Late chunking: embed first, split second
A newer trick worth knowing because it attacks the orphaned-chunk problem head-on. With a long-context embedding model (Jina v3, Voyage), embed the entire document in one forward pass so every token attends to the whole document, then mean-pool the token embeddings into per-chunk vectors. Each chunk vector now carries document-level context it could never get if embedded in isolation. Reported gains: +3.6% relative on BeIR and +24% relative on LongEmbed with 512-token chunks. The failure mode: it hurts when a document’s tail is unrelated to its head — pooling drags in irrelevant context — so it’s for coherent documents, not stapled-together ones.
The clean way to see why late chunking helps: in naive chunking, the token “it” in chunk 4 was embedded with no idea that “it” meant “the Q2 revenue figure” introduced in chunk 1 — the reference is severed at split time. In late chunking, every token attends to the whole document before the split, so “it” is already disambiguated inside its vector. It’s the same intuition as contextual retrieval (next), but achieved by the encoder’s attention rather than by an LLM writing an explicit prefix. Interview angle. If asked to compare them: late chunking is cheaper (no extra LLM call, one forward pass) but needs a long-context encoder and coherent docs; contextual retrieval is pricier but model-agnostic, works on any encoder, and also feeds the BM25 index. Naming that distinction is a strong signal.
Contextual retrieval — Anthropic’s fix for the orphaned chunk
Here’s the single most common chunk-level failure, made concrete. A chunk reads: “The company’s revenue grew 3% over the previous quarter.” Retrieved on its own, the model has no idea which company or which quarter — and a query like “ACME Q2 2023 revenue” may not even rank it. The chunk is locally fluent and globally useless.
Anthropic’s Contextual Retrieval (Sept 2024) fixes this at index time: before embedding each chunk, a cheap model (Claude Haiku) is handed the whole document and writes a 50–100 token context that situates the chunk — e.g. “From ACME Corp’s Q2 2023 10-Q; ‘revenue’ refers to ACME’s Q2 2023 total revenue vs Q1 2023.” That context is prepended to the chunk, and the contextualized chunk is indexed in both the embedding index and the BM25 index. The impact on top-20 retrieval failure rate stacks cleanly:
1# Contextual retrieval at index time: situate each chunk in its document.2# Prompt-cache the WHOLE document once, then reuse it for every chunk -> ~90% cheaper.3CONTEXT_PROMPT = (4 "Here is the whole document:\n{doc}\n\n"5 "Here is a chunk we want to situate within it:\n{chunk}\n\n"6 "Give a short (50-100 token) context that situates this chunk in the document "7 "for search -- resolve names, dates, and references. Answer only with the context."8)910def contextualize(doc, chunks, llm):11 out = []12 for c in chunks:13 ctx = llm(CONTEXT_PROMPT.format(doc=doc, chunk=c), cache_prefix=doc) # doc cached14 out.append(ctx + "\n\n" + c) # prepend, then index in BOTH embeddings + BM2515 return out16# Idempotent: re-run on a document UPDATE, or old and new chunk reps mix and shear recall.
The economics are why this is practical, not academic: contextualizing is a one-time index cost of ~$1.02 per million document tokens, and because you re-read the same document for all of its chunks, prompt caching cuts that by up to ~90%. Senior gotchas: the lift is largest with strong embedding models (Gemini, Voyage) and tapers past the top-20; and the contextualization step must be idempotent — re-run it on a document update or you mix old and new representations and silently shear recall (see scale, below).
Interview angle. Contextual retrieval is the single best “show me you read past the README” story in RAG, because the numbers are crisp and stackable: 5.7% → 3.7% (contextual embeddings) → 2.9% (+ contextual BM25) → 1.9% (+ reranker), a 67% cut in top-20 failures with no model change. When an interviewer says “your chunks lose their context, what do you do?”, walk that ladder and quote the one-time cost (~$1.02 / M tokens, ~90% off with caching). It signals you optimize the index, not just the prompt — and that you know the cheap wins compound.
Cursor chunks the hardest corpus there is — code that changes every keystroke, with a brutal long tail (most files are queried rarely) and hotspots (a few are queried constantly). It doesn’t split by token count; it splits on language-aware boundaries (function, class, import block), embeds at parse time, keys each vector by file path, and re-syncs the index continuously rather than on a schedule — namespacing by repo so one codebase can never retrieve into another. The generalizable lesson: chunk on semantic boundaries, embed-and-store at ingest, re-sync incrementally.
Why code breaks token-count chunking specifically: a function’s meaning lives in its signature and body together, a class method needs its class name, and a call needs its import — split any of those and the chunk is locally fluent but unanswerable for “where is X defined / who calls Y.” This is the same orphaned-chunk problem as prose, just sharper, which is why AST/tree-sitter splitting (split at function and class nodes, attach the file path and enclosing symbol as metadata) is the right tool. The transferable principle for any structured corpus — contracts, API docs, runbooks — is: split on the document’s own semantic units, not a token ruler, and carry the unit’s identity as metadata.
The scaling trap hiding inside chunking is the chunking clock. Change your chunker — new boundaries, new size, new contextualization prompt — and every chunk now represents different text, so its embedding is no longer comparable to the old ones (“representation shearing”). Re-chunking therefore always means a full re-embed; mixing generations in one index quietly collapses recall even though every individual vector looks healthy. At ~1 TB corpora a full re-embed runs on the order of $12k/month, which is why teams re-chunk behind a dual index and cut over only once the new index validates on a held-out set.
The failure here is uniquely nasty because it’s invisible. No request errors. Latency and throughput are healthy. Every individual vector looks fine. But because half the index represents old text and half represents new, the distribution of distances has shifted and cross-generation comparisons are subtly wrong, so recall sags a few points and stays there — and you only find out from a user complaint weeks later, if at all. Interview angle. “How do you safely change your chunking strategy in production?” → never in place; build the new index alongside the old (dual index), validate it on a held-out labelled set, then cut traffic over atomically and delete the old one. Saying “in place” here is a red flag that you haven’t run RAG at scale.
Small-to-big: decouple what you match from what you read
The clean way to get precision and context: index small units (256 tokens, or even single sentences) for sharp matching, but at generation time return the larger parent (≈1024 tokens) the match came from, capping the assembled context around 8k tokens. This “parent-document / small-to-big / sentence-window” pattern reports 5–15% end-to-end gains on entity-focused queries. Failure modes to guard: parent expansion drags in low-signal neighbours the model then hallucinates from, and several small hits inside one parent return duplicate context — so dedupe by parent id, keeping the highest child score.
python
1# index children (small) -> retrieve children -> return de-duped parents2def retrieve_small_to_big(query, k_children=20, k_parents=5):3 children = vector_search(query, k=k_children) # 256-token units4 best_by_parent = {}5 for c in children:6 p = c["parent_id"]7 if p not in best_by_parent or c["score"] > best_by_parent[p]["score"]:8 best_by_parent[p] = c # dedupe: keep top child9 ranked = sorted(best_by_parent.values(), key=lambda c: c["score"], reverse=True)10 return [load_parent(c["parent_id"]) for c in ranked[:k_parents]] # ~1024-token context
Metadata is part of the chunk, not an afterthought
Every chunk should carry structured metadata, and the decision of what to attach is as load-bearing as the chunk text. Three roles it plays. (1) Filtering — fields like doc_type, date, tenant_id, and especially acl let you constrain retrieval to the legal/relevant subset before similarity even runs (this is the security boundary in L5: filter then retrieve). (2) Citation — doc_id, title, section, and a stable URL/anchor are what turn a grounded answer into a citable one; without them the model can’t point at its source. (3) Freshness — a last_verified timestamp lets you time-filter stale content and powers the drift telemetry in L5. The senior framing: the chunk is the text plus the metadata that makes it findable, filterable, citable, and expirable.
Interview angle. A subtle but high-signal point: structured metadata filtering (a date range, a tenant id, a document type) is usually better handled by a metadata predicate on the query than by hoping the embedding captures it — embeddings are poor at exact constraints like “only 2024 docs.” Saying “I’d push hard filters into metadata and let the vector handle semantics” shows you know the division of labour, and it sets up the filtered-ANN gotcha (a filter that excludes most candidates can wreck recall — L5) that interviewers probe next.
Semantic chunking: how it works, and why it’s often not worth it
Semantic chunking deserves a precise mechanism because it’s frequently over-reached for. You embed each sentence, walk the document, and start a new chunk wherever the cosine distance between consecutive sentences spikes above a threshold — the idea being that a topic shift is a natural boundary. It’s appealing and sometimes helps on long, loosely-structured prose. But it adds an embedding pass at ingest, introduces a sensitive threshold to tune, and brings a real failure mode: it will happily cut between a question and its answer when the surface wording shifts, and it over-segments lists and dialogue. The practitioner verdict (Greg Kamradt’s “5 levels”, echoed across teams): reach for semantic chunking only after you’ve measured that boundary placement — not size, not context, not ranking — is your bottleneck. Recursive structure-aware splitting wins most of the time for a fraction of the cost.
python
1# Don't ship a chunking change on vibes -- sweep candidates against a labelled set.2def pick_chunk_size(eval_set, build_index, retrieve, sizes=(256, 384, 512, 768)):3 results = {}4 for size in sizes:5 build_index(chunk_size=size, overlap=int(size * 0.12)) # rebuild for each6 r20 = recall_at_k(eval_set, retrieve, k=20) # did we fetch it?7 p5 = precision_at_k(eval_set, retrieve, k=5) # is top-5 clean?8 results[size] = (round(r20, 3), round(p5, 3))9 return results10# Read it as a tradeoff curve: pick the size that holds recall@20 while maximizing11# precision@5. If the best p5 is still < 0.7, that's your signal to add a reranker.
A retrieved chunk says “revenue grew 3% over the previous quarter,” but the model can’t tell which company or quarter — and the right chunk often doesn’t rank for “ACME Q2 2023 revenue.” Strongest fix?
AUse much larger chunks so each one carries more surrounding textBContextual retrieval: prepend an LLM-written context to each chunk before indexing it (embeddings + BM25)CSwitch to a larger generation model
Your RAG has healthy recall@20, but answers are vague and miss specifics; precision@5 measures 0.55. Best first move?
AIncrease chunk size to give each chunk more contextBShrink chunks toward ~256 tokens and add a reranker to sharpen the top-5CSwap the embedding model for a larger one
You need to change your chunk size and contextualization prompt on a live 1 TB index. What’s the safe rollout?
ARe-chunk and re-embed in place, one document at a time, during low trafficBBuild a second index with the new chunker, validate it on a held-out labelled set, then cut traffic over atomicallyCKeep both chunkers’ vectors in one index and let RRF sort it out
Indexing a codebase, answers to “where is parseConfig defined?” keep missing or returning fragments. Your splitter is fixed-size 512 tokens. Best fix?
ASplit on language-aware boundaries (functions/classes via AST), attaching file path + enclosing symbol as metadataBIncrease the chunk size to 2,048 tokens so more code fitsCAdd overlap so fragments are duplicated across chunks
Your corpus is mostly scanned PDFs with tables. Recall is poor and answers garble numeric facts. Where do you look first?
ASwap to a semantic chunkerBRaise k and add a rerankerCThe parsing stage — use layout/OCR-aware extraction and keep tables as Markdown/HTML before chunking
Chunking questions test tradeoff fluency, not memorized numbers. Interviewers want to hear: you size by metric (precision@5), you separate the match-unit from the read-unit, you know the orphaned-chunk failure and its two fixes (late chunking / contextual retrieval), you treat parsing as its own stage, and you can change chunking in production without shearing recall. Lead with the tension, then the lever and the number.
01“What chunk size and why?” → start ~512 with structure-aware splitting, then measure precision@5 and drop toward ~256 if it sags — never a fixed answer.
02“Why don’t bigger chunks just give more context?” → one pooled vector dilutes the relevant signal across irrelevant tokens, similarity falls, the chunk ranks lower.
03“A chunk says ‘revenue grew 3%’ with no company/quarter — fix?” → contextual retrieval: LLM-written context prepended before indexing (embeddings + BM25); ~67% fewer top-20 failures with a reranker.
04“Late chunking vs contextual retrieval?” → late chunking = one long-context forward pass, cheap, needs coherent docs; contextual = extra LLM call, model-agnostic, also feeds BM25.
05“How do you get precision AND context?” → small-to-big: index ~256-token children for matching, return the ~1024-token parent to read, dedupe by parent id.
06“How do you chunk code / contracts / API docs?” → split on the document’s own units (AST functions, clauses, headings), carry the unit’s identity as metadata, not a token ruler.
07“Change chunking on a live index — how?” → dual index, validate on a held-out set, atomic cutover; in place shears recall silently.
08“Where does chunking quality actually get lost first?” → parsing — tables/scans/multi-column PDFs corrupt text before any chunker runs.
Going deeper, the strong-candidate follow-ups: “how do you know your chunking is the problem and not the generator?” (split the eval — recall@k vs precision@5 vs faithfulness, L6); “your precision is fine but recall is low” (chunks too small / context stripped, or an ANN-param ceiling, L1 — not a chunk-size-down move); “parent expansion is hurting answers” (it dragged in low-signal neighbours the model hallucinated from — cap context and dedupe by parent); and “is contextual retrieval worth the cost?” (yes when chunks lose identifiers and you have strong embeddings; the one-time cost is cache-cheap, the lift tapers past top-20). Always tie the symptom to a metric.
Could you pick a chunking + contextualization strategy for a messy 50k-doc corpus and defend the chunk size with a metric?
Not yetMostlyConfident
Takeaways
Parsing is stage zero — tables/scans/multi-column PDFs corrupt text before any chunker runs.
Match unit ≠ meaning unit: small chunks to match, parent expansion to read (dedupe by parent).
Don’t guess chunk size — measure precision@5. Start ~512, drop toward ~256 if it sags. Bigger dilutes.
Orphaned chunks: fix with late chunking (cheap, coherent docs) or contextual retrieval (model-agnostic, feeds BM25).
Contextual retrieval cuts top-20 retrieval failure ~49% (≈67% with a reranker) for a one-time, cache-cheap pass.
Re-chunking = full re-embed; cut over behind a dual index or you shear recall silently.
Next: hybrid retrieval and the multi-stage ranking pipeline that systems like Perplexity actually run at 200M queries a day.