Lesson 1 of 7 · 46 min

Retrieval foundations & the shapes of RAG

Why retrieval — not the model — caps your whole system, how embeddings and ANN actually work, the embedding choices that move the needle, and the three production shapes of RAG you must pick between before you write a line.

The thing that actually caps quality

A senior mental model up front: in RAG, the model is rarely the bottleneck — retrieval is. If the answer-bearing chunk never makes the top-k, no model, however large, can ground an answer on it. Most production “hallucinations” are retrieval misses wearing a costume. So we debug retrieval first, and we measure it like the information-retrieval problem it is. This lesson is the ground floor for the whole track — and the part of a system-design interview where most candidates wave their hands.
Why does a vanilla LLM hallucinate on your internal data at all? Because the weights are a lossy compression of the public web at training time — they have no row for your Q3 board deck, your incident runbook, or yesterday’s ticket. RAG is the fix that won: instead of fine-tuning the facts into the model (expensive, stale the moment a doc changes, hard to attribute), you retrieve the relevant text at query time and put it in the context. That buys you three things fine-tuning can’t — freshness (re-index, don’t re-train), attribution (cite the chunk), and access control (filter by who may see what). Interviewers love “RAG vs fine-tuning” precisely because the naive answer (“fine-tune on our docs”) misses all three.
Keyword search matches strings — ask for “car” and it never finds a doc that only says “automobile.” Semantic search matches meaning, and that shift is what makes RAG work. The mechanism is the embedding: a model maps text to a fixed-length vector (say 1,536 numbers) so that similar meaning lands at nearby points. To search, you embed the query and take the document vectors with the highest cosine similarity.
A few mechanics worth saying precisely, because interviewers probe them. The embedding model is a transformer encoder that pools token vectors into one fixed-length vector — usually mean-pooling or a special CLS token — so a 3-word query and a 400-word chunk land in the same dimensional space and are directly comparable. We use cosine similarity (angle), not Euclidean distance, because document length inflates raw magnitude and we care about direction of meaning, not length. With normalized vectors, cosine and dot-product rank identically, which is why most vector DBs default to one of those two metrics.
python
1from openai import OpenAI2import numpy as np34client = OpenAI()56def embed(texts):7    r = client.embeddings.create(model="text-embedding-3-small", input=texts)8    return np.array([d.embedding for d in r.data])910docs = ["The cheetah is the fastest land animal.",11        "Postgres is a relational database.",12        "Sprinters train explosive speed."]13doc_vecs = embed(docs)1415def search(query, k=2):16    q = embed([query])[0]17    sims = doc_vecs @ q / (np.linalg.norm(doc_vecs, axis=1) * np.linalg.norm(q))18    return [docs[i] for i in sims.argsort()[::-1][:k]]1920search("what runs quickly?")   # -> the speed docs, without the word "quick"
Interview angle. “Why cosine and not Euclidean?” → cosine ignores magnitude, so a long doc and a short query compare on meaning, not length; and on L2-normalized vectors cosine, dot-product, and Euclidean all induce the same ranking, so the choice is mostly about whether your index normalizes. Being able to say that — and that the answer is a sample-free, deterministic distance, unlike the generator — signals you actually understand the retrieval layer.

Why ANN, not brute force — the index is the engine

At scale you never compare the query against every vector. Exact (“flat”) search is O(N·d): a single query against 100M 1,536-d vectors is ~150 billion multiply-adds — tens of seconds on a CPU, far past any interactive budget. A vector database instead builds an approximate-nearest-neighbour (ANN) index that trades a few points of recall for 100–1000× speed. The pipeline is always embed query → ANN search → top-k → (rerank) → LLM.
The dominant index is HNSW (Hierarchical Navigable Small World): a multi-layer proximity graph where you greedily hop from a sparse top layer down to dense lower layers, each hop moving closer to the query. Two knobs decide its quality/cost: M (edges per node — higher = better recall, more RAM) and efSearch (candidates explored per query — higher = better recall, more latency). The senior framing: HNSW is a recall/latency/memory triangle you tune, not a black box. Alternatives — IVF (cluster, then probe a few clusters) and IVFPQ (IVF plus product quantization that compresses each vector to a handful of bytes) — trade recall for a ~20× memory cut. We give the at-scale economics of all three their own lesson (L5); here, just know the index — not the embedding model — is usually what sets your p99 and your RAM bill.
code
1Search cost as the corpus grows (1,536-d vectors, rough orders of magnitude)23  Method     1M vectors        100M vectors      Recall    Memory4  --------   --------------    --------------    ------    ------------------5  Flat/exact ~50-200 ms        seconds+          100%      raw vectors only6  HNSW       ~1-5 ms           ~5-30 ms          0.95-0.99 1.5-2x raw (graph)7  IVF        ~2-8 ms           ~10-40 ms         0.90-0.97 ~1.1x raw8  IVFPQ      ~3-10 ms          ~20-60 ms         0.80-0.92 ~0.05x raw (PQ)910  Takeaway: ANN buys 100-1000x speed for a few points of recall.11  Index choice sets p99 and RAM far more than the embedding model does.

The three shapes of RAG — pick one before you build

The most common senior mistake is building “a RAG” without deciding which shape of RAG you have. The shape sets your latency budget, your index footprint, and which lever matters most. Across documented production systems, they cluster into three:
  1. 01Chat-shaped (Notion, DoorDash, Uber Genie) — answer over a closed corpus; 300 ms–2 s budget; modest QPS. Dominant lever: chunking + reranking + model choice, not the DB.
  2. 02Enterprise-search (Glean, Elastic, Harvey) — search-then-summarize across a corporate corpus; sub-second p99; multi-tenancy + RBAC dominate. Dominant lever: the hybrid index.
  3. 03Web-shaped (Perplexity, LinkedIn feed) — ask the open web / a billion-doc index; freshness in seconds; tight latency. Dominant lever: one engine fusing lexical + dense + freshness signals.
Perplexity’s search runs ~200M queries/day at p50 358 ms; Notion cut chat latency from ~2 s to ~350 ms by swapping to a smaller fine-tuned model — not by touching retrieval. Same word, “RAG,” three completely different engineering problems. Teams that “failed at RAG” usually picked the wrong shape.
The shape also dictates your freshness SLA and ingestion cadence, which quietly drive cost. Chat-shaped corpora change daily or weekly, so a nightly batch re-index is fine and cheap. Enterprise-search corpora change as employees edit docs, so you need change-data-capture (CDC) ingestion and per-tenant isolation. Web-shaped systems demand seconds-fresh updates — Perplexity sustains tens of thousands of index updates per second — which is why they build on a single co-located engine (Vespa) rather than stitching a vector DB to a separate crawler. Interview angle. When asked to “design a RAG,” your first move should be to name the shape and its freshness SLA out loud; it instantly frames every later decision and shows you’ve built one, not just read about it.
The word “RAG” hides three different systems. Most “we tried RAG and it didn’t work” stories are really “we built the chat shape’s architecture for an enterprise-search problem.” Name the shape first.

Why naive top-k breaks: the seven failure points

Barnett et al. catalogued seven failure points that recur in real RAG systems. Knowing them turns “the bot is dumb” into a precise diagnosis — and each maps to a fix in a later lesson. This taxonomy is gold in an interview: instead of “I’d improve the prompt,” you can localize a failure to a stage and name the lever.
code
11. Missing content      — the answer isn't in the corpus            -> ingestion / coverage22. Missed top-ranked    — it's retrieved, but below your cutoff       -> reranking (L4)33. Not in context       — right doc, wrong chunk (split/buried)       -> chunking + contextual (L2)44. Not extracted        — it's in context, the model misses it        -> prompt / grounding (L7)55. Wrong format         — ignored your "return JSON" etc.             -> structured output66. Wrong specificity    — too vague or too narrow                     -> query transformation (L4)77. Incomplete           — partial answer to a multi-part question     -> decomposition (L4)
Notice the split: failures 1–3 are retrieval problems (the right text never reached the prompt) and 4–7 are generation/orchestration problems (the text was there, the model or pipeline mishandled it). The diagnostic discipline is to locate the boundary first: if the gold chunk is in the retrieved set but the answer is wrong, no amount of chunking or reranking helps — you’re in generation territory (grounding prompt, L7). If it’s not in the set, no prompt engineering helps. Most teams burn weeks tuning the wrong half because they never instrumented the boundary.
Make those two metrics concrete, because interviewers ask you to define them. Recall@k = fraction of queries whose gold chunk appears in the top-k; it answers “did we even fetch the answer?” and it’s the metric retrieval must maximize, because a miss here is unrecoverable downstream. Precision@k = fraction of the top-k that is relevant; low precision means you’re packing the context with noise, which both costs tokens and degrades the generator (more context, more “lost in the middle”). The standard split: tune your first stage for recall@50–100, then let a reranker (L4) buy precision@5. When you can only quote one number, quote recall@k for the retriever — it’s the ceiling on everything else.
python
1# Define the two metrics every RAG eval starts with -- on a labelled set.2def recall_at_k(eval_set, retrieve, k=20):3    hits = 04    for query, gold_chunk_ids in eval_set:5        got = {r["chunk_id"] for r in retrieve(query, k=k)}6        hits += 1 if got & set(gold_chunk_ids) else 0   # did ANY gold chunk make top-k?7    return hits / len(eval_set)89def precision_at_k(eval_set, retrieve, k=5):10    total = 0.011    for query, gold_chunk_ids in eval_set:12        got = [r["chunk_id"] for r in retrieve(query, k=k)]13        total += len(set(got) & set(gold_chunk_ids)) / k   # how much of top-k was relevant?14    return total / len(eval_set)15# Rule of thumb: maximize recall@50-100 in stage 1, buy precision@5 with a reranker.

The labelled set is the prerequisite for everything

You cannot debug — or even claim to improve — retrieval without a small labelled set, and building one is the unglamorous first task every serious RAG team does. The minimum viable artifact is ~50–200 (query, gold-chunk-ids) pairs that look like real traffic. Three ways to get them, in increasing quality: (1) harvest real user queries from logs and label which chunk(s) answer each (best — it’s your true distribution); (2) have a strong LLM read each chunk and write a question it answers, then verify a sample by hand (fast, but watch contamination and skew); (3) hand-write the hard, edge-case, and adversarial queries you know will break it. The non-negotiable, foreshadowing L6: harvest the gold contexts from your production retriever’s runs, not from chunks you hand-picked, or your offline recall will read far higher than reality.
python
1# Bootstrap a labelled set from your OWN chunks (then hand-verify a sample).2def synth_eval_set(chunks, llm, n_per_chunk=1):3    rows = []4    for c in chunks:5        q = llm("Write one realistic question this passage answers. "6                "Question only:\n" + c["text"])7        rows.append({"query": q, "gold_chunk_ids": [c["chunk_id"]]})8    return rows                       # then: human-verify ~20%, prune leaky/duplicate Qs910# Use it to compare ANYTHING -- embedding model, chunk size, hybrid on/off:11baseline = recall_at_k(eval_rows, retrieve_v1, k=20)12candidate = recall_at_k(eval_rows, retrieve_v2, k=20)13# Ship the change only if candidate beats baseline on YOUR distribution.
Retrieval you can’t measure is retrieval you can’t improve. The first hour of any RAG project is better spent building 50 labelled queries than swapping the embedding model — every later decision is an A/B test against that set.

Choosing an embedding model

Decide on four axes, in order: domain fit (does it understand your jargon? test on your data, not MTEB), dimension (bigger = more RAM and slower ANN; many models support Matryoshka truncation to trade a little recall for a lot of memory), context length (matters for late chunking, next lesson), and cost/latency at your QPS. MTEB is a starting filter, not a verdict — leaderboard rank rarely survives contact with a specialized corpus.
Two practical constraints juniors miss. First, the asymmetry of query vs document: some models (and most retrieval-tuned ones — e5, BGE, Nomic) want a task prefix like “query:” / “passage:”, and getting it wrong silently halves recall because queries and docs land in slightly different sub-spaces. Test that you’re prefixing correctly before you blame the model. Second, you are married to your embedding model: changing it invalidates every stored vector (a new model’s space is not comparable to the old one), so a swap means a full, paid re-embed of the entire corpus. That lock-in — covered as the “embedding clock” in L5 — is why “just try a bigger encoder” is rarely the cheap experiment it sounds like.
Matryoshka embeddings (the OpenAI v3 family, Nomic, many 2024+ models) are trained so the first N dimensions are independently meaningful — you can truncate 3,072-d down to 512-d or 256-d and keep most of the recall while cutting RAM and ANN latency several-fold. The senior move at scale: store a truncated vector for the fast first-stage ANN search, then optionally re-score with the full vector or a reranker. Interview angle. “Your index RAM is too high — what do you do without re-embedding?” → truncate Matryoshka dimensions and/or switch to product quantization (IVFPQ, L5); both shrink the footprint at a measured recall cost you recover with reranking.
code
1Picking an embedding model -- the axes that actually decide it23  Axis            What to check                       Failure if ignored4  ----------      -------------------------------     ---------------------------5  Domain fit      eval on YOUR corpus, not MTEB        leaderboard win, prod loss6  Dimension       RAM = N_vectors x dim x 4 bytes      20x RAM surprise at scale7  Query prefix    "query:"/"passage:" if required      ~halved recall, silent8  Context length  >= your chunk size (+ late chunk)    truncated chunks, lost text9  Cost / QPS      $/1M tokens x ingest + query rate    embedding bill > LLM bill10  Lock-in         swap = full re-embed                 "cheap experiment" isn't1112  RAM example: 50M chunks x 1,536 dim x 4 bytes = ~300 GB raw (before the index).
Notice that last RAM line — it surprises people. Fifty million 1,536-d float32 vectors are ~300 GB of raw data before HNSW adds its graph (1.5–2×). That single arithmetic fact is why dimension is a first-order cost decision, not a footnote, and why Matryoshka truncation and quantization exist. Being able to do this multiplication on a whiteboard is a reliable senior signal.
articlePatterns for Building LLM-based Systems & ProductsEugene YanpaperMatryoshka Representation Learning (truncatable embeddings)Kusupati et al. (arXiv)docsMTEB — Massive Text Embedding Benchmark leaderboardHugging FacerepoRAG_Techniques — simple RAG, embeddings & evaluation notebooksNirDiamant/RAG_Techniques

Checkpoint

A teammate proposes fixing weak answers by upgrading to the largest available embedding model. What’s the senior response?

AAgree — bigger embeddings are the main driver of retrieval qualityBFirst measure recall@k / precision@k to locate the failure, then fix chunking/hybrid before swapping modelsCSwitch to a larger generation model instead
Sign up free to answer and see why

Checkpoint

You’re building support QA over a fixed 40k-doc knowledge base with a 1.5 s latency budget. Which “shape” framing should drive your architecture?

AWeb-shaped — optimize a single engine for seconds-fresh, billion-doc searchBChat-shaped — closed corpus, comfortable latency; spend your effort on chunking, reranking, and model choiceCIt doesn’t matter — all RAG is the same pipeline
Sign up free to answer and see why

Checkpoint

In an interview you’re asked “RAG or fine-tuning to make the model answer from our internal wiki?” Strongest opening?

AFine-tune the model on the wiki so the facts live in the weightsBRAG — retrieve at query time for freshness, citations, and access control; consider fine-tuning only for style/format or a fixed skillCNeither — just use a bigger base model
Sign up free to answer and see why

Checkpoint

Your retriever’s recall@20 is stuck at 0.86 and tuning chunking hasn’t moved it. You’re on HNSW with default params. What’s worth trying first?

ARaise efSearch (and/or M) — ANN is approximate, so the index config may be the recall ceilingBImmediately swap to a larger embedding modelCLower k to 5 so precision improves
Sign up free to answer and see why

Checkpoint

You’re serving 50M chunks and the vector index won’t fit in your RAM budget. Which move shrinks the footprint WITHOUT re-embedding the corpus?

ARe-embed everything with a smaller modelBIncrease the chunk size so there are fewer vectorsCTruncate Matryoshka dimensions and/or switch to product quantization (IVFPQ), recovering quality with a reranker
Sign up free to answer and see why

Interview prep

Foundations rounds for RAG roles test whether you treat retrieval as an information-retrieval system with real metrics and tradeoffs — or as a magic box. Interviewers probe four things: do you debug retrieval before generation, can you define recall@k / precision@k, do you understand embeddings + ANN mechanically, and can you size the thing (RAM, cost, latency) on a whiteboard. Lead with the mechanism, then the production implication.
  1. 01“What actually caps a RAG system’s quality?” → retrieval, not the model — a chunk that never makes top-k can’t be grounded; most “hallucinations” are retrieval misses.
  2. 02“RAG vs fine-tuning?” → RAG for volatile, access-controlled facts (freshness, citations, ACL); fine-tuning for style/format or a fixed skill — they compose, not compete.
  3. 03“Why cosine, not Euclidean?” → cosine ignores magnitude so length doesn’t distort meaning; on normalized vectors cosine/dot/L2 rank identically.
  4. 04“Why ANN instead of comparing every vector?” → exact search is O(N·d) and seconds-slow at 100M; ANN buys 100–1000× speed for a few points of recall.
  5. 05“How does HNSW work and what knobs matter?” → greedy descent through a layered proximity graph; M (edges) and efSearch (candidates) trade recall for RAM/latency.
  6. 06“Define recall@k vs precision@k.” → recall@k = did the gold chunk make top-k (the ceiling); precision@k = how much of top-k was relevant (noise/cost).
  7. 07“Estimate the index RAM for 50M chunks.” → N × dim × 4 bytes ≈ 300 GB raw at 1,536-d, × ~1.5–2 for HNSW — then mention Matryoshka/PQ to cut it.
  8. 08“Three engineers each say ‘we built a RAG’ — what do you ask?” → which shape (chat / enterprise-search / web), the freshness SLA, and the QPS; they imply totally different architectures.
To go deeper, expect the follow-ups that separate “read a blog” from “shipped one”: “walk me through the failure when answers are wrong — how do you localize it?” (instrument the retrieval/generation boundary: is the gold chunk in the retrieved set or not?); “your recall is fine but answers are noisy” (precision problem → reranker, L4; or chunks too big, L2); “the embedding bill is bigger than the LLM bill” (batch ingestion, cache embeddings of unchanged docs, truncate dims); and “you swapped encoders and recall dropped” (representation shearing / wrong query prefix — L5). In every case, name the metric and the stage before proposing a fix.

Could you state which RAG shape a new project is, and name the metric you’d use to debug its retrieval?

New to itGetting thereConfident

Takeaways

  • Retrieval caps the system — debug it first, and measure recall@k (the ceiling) / precision@k.
  • RAG beats fine-tuning for volatile, access-controlled facts: freshness, citations, ACL.
  • ANN (HNSW) is approximate — its config (M, efSearch) is a recall ceiling you tune, not the encoder.
  • Pick a shape (chat / enterprise-search / web) and name its freshness SLA before you build.
  • Embedding choice is domain-fit first, size last; dimension is a RAM bill (N × dim × 4 bytes).
  • The seven failure points turn “it’s dumb” into a precise, stage-localized diagnosis.

Next: how you split documents — and the contextual-retrieval trick that cuts retrieval failure by half.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.