Lesson 3 of 7 · 44 min
Hybrid retrieval & the multi-stage pipeline
Dense misses exact tokens; BM25 misses paraphrase. How BM25 actually scores, why you fuse with RRF, the multi-stage pipeline Perplexity and Uber run in production, sparse-neural retrievers, and when hybrid isn’t worth it.
BM25) nails those but misses paraphrase. They fail on different queries, which is exactly why fusing them wins: hybrid retrieval runs both and combines the ranked lists.Why the “obsolete” technique keeps winning
How BM25 actually scores — and why it’s a strong baseline
k1), and long documents are discounted so they don’t win just by being long (controlled by b). That’s why BM25 nails exact, rare, discriminative tokens — and why on many real corpora a tuned BM25 ties or beats a generic dense model.“E-1042” into nothing matchable — and your exact-match advantage, the whole reason you added BM25, silently evaporates. Interview angle. “Hybrid is on but exact codes still don’t match” is almost always an analyzer bug, not a fusion bug; saying that distinguishes someone who has run a lexical index from someone who has only read about RRF.1BM25 vs dense -- which retriever wins, by query type23 Query example Dense BM25 Use hybrid because...4 ------------------------------- ----- ----- ---------------------5 "how do I cancel my plan" strong weak paraphrase -> dense6 "error E-1042" weak strong exact rare token -> BM257 "ACME-X200 firmware" weak strong SKU / part number -> BM258 "what's the refund window" strong medium intent -> dense, words -> both9 "Dr. Sarah Okonkwo paper" weak strong proper noun -> BM251011 Dense and BM25 fail on DIFFERENT queries -> fusing covers both. That's hybrid.Common mistake
“BM25 is a dumb keyword count.”
Fusion: RRF, and when to weight it
Σ 1/(k + rank) across retrievers (k ≈ 60). It uses ranks, not scores — which is the whole point, since BM25 scores are unbounded and cosine sits in [-1, 1], so averaging the raw numbers is meaningless. That’s why most vector DBs ship RRF as the default hybrid mode.1def rrf(rankings, k=60):2 scores = {}3 for ranking in rankings: # e.g. [bm25_ids, vector_ids]4 for rank, doc_id in enumerate(ranking):5 scores[doc_id] = scores.get(doc_id, 0) + 1 / (k + rank)6 return sorted(scores, key=scores.get, reverse=True)78hybrid = rrf([bm25_search(q, 50), vector_search(q, 50)])[:50] # feed a reranker nextα·norm(bm25) + (1-α)·dense) can beat RRF by a few NDCG points if you’re willing to calibrate and accept slightly worse p99 — most teams take RRF’s zero-tuning robustness.1Hybrid vs single retrievers (NDCG, representative public benchmarks)23 Setup WANDS BEIR (system)4 --------------------------- -------- -------------5 BM25 only 0.6983 43.46 Dense only 0.6953 ~ comparable7 Hybrid (RRF) 0.7497 52.68 Hybrid + cross-encoder higher higher (L4)910 Alone, BM25 and dense tie; FUSED they jump (+7.4% WANDS, +9 BEIR pts).11 The lift is largest where queries mix paraphrase with exact tokens.α is corpus- and query-distribution-specific, so the moment your documents or your users drift, your tuned weight is stale and silently mis-ranking — and nobody re-tunes it until something breaks. RRF has no weight to go stale. Interview angle. When asked “RRF or weighted fusion?”, the senior answer is “RRF by default for its zero-maintenance robustness; weighted only if I have a labelled set to tune on and a plan to re-tune as the corpus drifts — and even then the gain is usually a few NDCG points.” That framing shows you weigh maintenance cost, not just benchmark deltas.Key idea
Sparse-neural retrievers: the third option
The multi-stage pipeline companies actually run
1Production multi-stage retrieval (the Perplexity / Uber shape)23 query4 -> [recall] BM25 + dense ANN (cast a wide net: ~top-50-100)5 -> [fuse] RRF (merge the two ranked lists)6 -> [precision] cross-encoder rerank (score query x doc jointly: top-5)7 -> LLM (grounded generation + citations)Common mistake
“Vector search made keyword search obsolete.”
When hybrid isn’t worth it — and how it fails
- 01Skip it when queries are pure natural-language paraphrase over a small, clean corpus — you add latency and an index for no gain.
- 02Analyzer bug: a BM25 analyzer that lowercases and strips hyphens turns “E-1042” into nothing — your exact-match advantage evaporates. Tune the analyzer.
- 03RRF can be dominated by the noisier retriever if it returns far more candidates — cap each list (e.g. 50) and prune.
- 04k is a knob: k≈60 is the TREC default; with many relevant docs per query, k=10–30 can help.
The recall stage, end to end
1# Hybrid recall -> fuse -> hand off to the reranker (the Perplexity/Uber shape).2def hybrid_recall(query, k_each=50):3 bm25_ids = bm25_search(query, k=k_each) # exact tokens, codes, names4 dense_ids = vector_search(query, k=k_each) # paraphrase, intent5 fused = rrf([bm25_ids, dense_ids], k=60) # rank-based, no calibration6 return fused[:k_each] # capped shortlist for the reranker78def retrieve(query):9 shortlist = hybrid_recall(query, k_each=50) # RECALL: wide, cheap10 return rerank(query, shortlist, top_n=5) # PRECISION: costly, on shortlist11# Two stages, two jobs: hybrid maximizes the chance the answer is PRESENT;12# the reranker maximizes that the right 5 are at the TOP.Tuning BM25: k1 and b are not magic constants
Key idea
Checkpoint
A user searches for the exact error code “E1042.” Which retriever most reliably surfaces it?
Checkpoint
Why do production systems like Perplexity fuse BM25 + dense and then run a cross-encoder, instead of just doing one bigger vector search?
Checkpoint
You turned on hybrid, but searches for part numbers like “ACME-X200” still miss even though the docs contain them. Most likely cause?
Checkpoint
A teammate wants to replace RRF with a tuned weighted fusion (α·bm25 + (1−α)·dense) to squeeze a few NDCG points. What’s the senior caveat?
Checkpoint
Queries are pure natural-language paraphrase over a small, clean FAQ with no codes or proper nouns. A teammate insists on adding BM25 + RRF. Best call?
Interview prep
- 01“Did vector search kill keyword search?” → no — dense compresses away rare exact tokens; BM25 keeps them. They cover different queries, so you fuse.
- 02“How does BM25 score?” → IDF-weighted term matches with saturation (k1) and length normalization (b) — not a raw count; still a top BEIR baseline.
- 03“Why RRF and not averaging scores?” → BM25 is unbounded, cosine is [-1,1]; fuse on rank so neither scale dominates, and cap each list before fusing.
- 04“RRF vs weighted fusion?” → RRF by default (zero maintenance); weighted only with a labelled set and a re-tuning plan, for a few NDCG points.
- 05“Why fuse then rerank instead of one big vector search?” → recall (wide cheap net) and precision (costly cross-encoder on the shortlist) are separate stages.
- 06“What is SPLADE / learned sparse?” → transformer-predicted sparse weights with term expansion — BM25’s exactness plus some synonym handling, at a larger index.
- 07“Hybrid is on but codes don’t match — why?” → analyzer bug (it stripped the punctuation/hyphen); fix tokenization, not fusion.
- 08“When do you skip hybrid?” → pure paraphrase over a small clean corpus with no codes/names — measure; the extra index and latency may buy nothing.
Common mistake
The red flag that sinks candidates: treating hybrid as a magic “better retrieval” switch.
Could you justify a hybrid + multi-stage design — and when to skip hybrid — in a design review?
Takeaways
- Dense = meaning/recall; BM25 = exact-term precision (IDF + saturation + length-norm). Fuse with RRF (rank-based, no calibration).
- RRF beats weighted fusion on maintenance: no corpus-specific weight to go stale on drift.
- Learned sparse (SPLADE) adds term expansion to BM25 — complement, rarely replacement.
- Hybrid is the recall stage of a multi-stage ranker — Perplexity/Uber fuse, then rerank, on one co-located engine.
- Hybrid is near-free after chunking — but the analyzer is the usual failure, and skip it for paraphrase-only corpora.
- Pruning the document set (Uber Genie, −60% wrong advice) can beat swapping encoders.
Next: the precision stage — cross-encoder reranking, ColBERT, and query transformation.
Sources
Free to read · better with Enzo
Learn it with Enzo
Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.