Lesson 3 of 7 · 44 min

Hybrid retrieval & the multi-stage pipeline

Dense misses exact tokens; BM25 misses paraphrase. How BM25 actually scores, why you fuse with RRF, the multi-stage pipeline Perplexity and Uber run in production, sparse-neural retrievers, and when hybrid isn’t worth it.

Dense (vector) retrieval is great at meaning but surprisingly bad at exact tokens — error codes, SKUs, function names, rare acronyms, person names. Keyword search (BM25) nails those but misses paraphrase. They fail on different queries, which is exactly why fusing them wins: hybrid retrieval runs both and combines the ranked lists.

Why the “obsolete” technique keeps winning

Every few months someone declares keyword search dead — and every production retrieval team keeps BM25 in the stack. The reason is structural, not nostalgic: a dense vector compresses text into a few hundred numbers, and rare exact tokens — the SKU, the error code, the surname — are precisely the information a lossy compressor throws away. BM25 keeps them because it never compresses. Hybrid isn’t a hedge; it’s two retrievers covering each other’s blind spots.

How BM25 actually scores — and why it’s a strong baseline

BM25 is a bag-of-words ranking function with two ideas bolted on that make it shockingly hard to beat. (1) IDF (inverse document frequency): rare query terms count more — matching “pneumonoultramicroscopic” is worth far more than matching “the.” (2) Saturation + length normalization: the 10th occurrence of a term adds less than the 2nd (controlled by k1), and long documents are discounted so they don’t win just by being long (controlled by b). That’s why BM25 nails exact, rare, discriminative tokens — and why on many real corpora a tuned BM25 ties or beats a generic dense model.
The senior gotcha lives in the analyzer, not the formula. BM25 only matches tokens your analyzer produced, so its behaviour is decided by tokenization, lowercasing, stop-word removal, and stemming. A default analyzer that splits on punctuation and strips hyphens turns the error code “E-1042” into nothing matchable — and your exact-match advantage, the whole reason you added BM25, silently evaporates. Interview angle. “Hybrid is on but exact codes still don’t match” is almost always an analyzer bug, not a fusion bug; saying that distinguishes someone who has run a lexical index from someone who has only read about RRF.
code
1BM25 vs dense -- which retriever wins, by query type23  Query example                     Dense    BM25     Use hybrid because...4  -------------------------------   -----    -----    ---------------------5  "how do I cancel my plan"         strong   weak     paraphrase -> dense6  "error E-1042"                    weak     strong   exact rare token -> BM257  "ACME-X200 firmware"              weak     strong   SKU / part number -> BM258  "what's the refund window"        strong   medium   intent -> dense, words -> both9  "Dr. Sarah Okonkwo paper"         weak     strong   proper noun -> BM251011  Dense and BM25 fail on DIFFERENT queries -> fusing covers both. That's hybrid.

Fusion: RRF, and when to weight it

The standard fusion is Reciprocal Rank Fusion (RRF): each retriever ranks the candidates, and each doc scores Σ 1/(k + rank) across retrievers (k ≈ 60). It uses ranks, not scores — which is the whole point, since BM25 scores are unbounded and cosine sits in [-1, 1], so averaging the raw numbers is meaningless. That’s why most vector DBs ship RRF as the default hybrid mode.
python
1def rrf(rankings, k=60):2    scores = {}3    for ranking in rankings:               # e.g. [bm25_ids, vector_ids]4        for rank, doc_id in enumerate(ranking):5            scores[doc_id] = scores.get(doc_id, 0) + 1 / (k + rank)6    return sorted(scores, key=scores.get, reverse=True)78hybrid = rrf([bm25_search(q, 50), vector_search(q, 50)])[:50]   # feed a reranker next
How much does it buy? On the WANDS benchmark, BM25 (NDCG 0.6983) and pure vector (0.6953) are statistically indistinguishable alone — but hybrid RRF jumps to 0.7497 (+7.4%). Across BEIR, hybrid lifts system NDCG from 43.4 to 52.6. A tuned convex combination (α·norm(bm25) + (1-α)·dense) can beat RRF by a few NDCG points if you’re willing to calibrate and accept slightly worse p99 — most teams take RRF’s zero-tuning robustness.
code
1Hybrid vs single retrievers (NDCG, representative public benchmarks)23  Setup                         WANDS      BEIR (system)4  ---------------------------   --------   -------------5  BM25 only                     0.6983     43.46  Dense only                    0.6953     ~ comparable7  Hybrid (RRF)                  0.7497     52.68  Hybrid + cross-encoder        higher     higher (L4)910  Alone, BM25 and dense tie; FUSED they jump (+7.4% WANDS, +9 BEIR pts).11  The lift is largest where queries mix paraphrase with exact tokens.
The deeper reason RRF wins on operations, not just quality: a learned weight α is corpus- and query-distribution-specific, so the moment your documents or your users drift, your tuned weight is stale and silently mis-ranking — and nobody re-tunes it until something breaks. RRF has no weight to go stale. Interview angle. When asked “RRF or weighted fusion?”, the senior answer is “RRF by default for its zero-maintenance robustness; weighted only if I have a labelled set to tune on and a plan to re-tune as the corpus drifts — and even then the gain is usually a few NDCG points.” That framing shows you weigh maintenance cost, not just benchmark deltas.

Sparse-neural retrievers: the third option

There’s a middle path worth knowing because interviewers increasingly ask about it: learned sparse retrievers like SPLADE. They produce a high-dimensional sparse vector (one weight per vocabulary term, mostly zeros) but the weights are predicted by a transformer and include term expansion — the document “car” gets non-zero weight on “automobile,” “vehicle,” “sedan.” So they keep BM25’s exact-match strength and inverted-index speed while adding some of dense retrieval’s synonym handling. The cost: larger indexes than BM25, a model in the ingest path, and expansion that can over-fire on ambiguous terms. The honest verdict: SPLADE can outperform raw BM25 and sometimes a generic dense model, but in production it usually complements rather than replaces a dense+BM25 hybrid — reach for it when exact-match matters but pure BM25’s vocabulary mismatch is hurting recall.

The multi-stage pipeline companies actually run

In production, hybrid is rarely the whole story — it’s the recall stage of a multi-stage ranker. Perplexity runs exactly this at ~200M queries/day on Vespa: dense embeddings + sparse lexical for recall, fed into a cross-encoder reranker for precision, with information and compute co-located on the same nodes to avoid the overhead of stitching a separate vector DB to a separate keyword engine. The result: p50 358 ms, p95 763 ms.
code
1Production multi-stage retrieval (the Perplexity / Uber shape)23  query4    -> [recall]   BM25  +  dense ANN        (cast a wide net: ~top-50-100)5    -> [fuse]     RRF                        (merge the two ranked lists)6    -> [precision] cross-encoder rerank      (score query x doc jointly: top-5)7    -> LLM (grounded generation + citations)
The architectural point in Perplexity’s design is the co-location: putting lexical, dense, and the documents themselves on the same Vespa nodes avoids the network hop and consistency headache of stitching a standalone vector DB to a separate keyword engine to a separate store. At 200M queries/day with tens of thousands of index updates per second, that stitching is where p99 and freshness go to die. The transferable lesson even at smaller scale: prefer one engine that does hybrid natively over three you have to keep in sync.
Uber’s Genie wraps the same hybrid core (dense + BM25 over enriched metadata) in agents — a Query Optimizer that rewrites/decomposes, a Source Identifier that narrows to the right document subset, and a Post-Processor that de-dupes and restores order. On a 100-query golden set it delivered +27% acceptable answers and −60% incorrect advice. The lesson: at support-grade accuracy, the failure is more often “wrong document set” than “wrong similarity function” — pruning before retrieval recovers more than swapping encoders.
That −60% incorrect-advice number deserves emphasis because it inverts a common instinct. Genie didn’t get safer by improving similarity; it got safer by narrowing what could be retrieved at all (Source Identifier) so the model never saw a plausibly-relevant-but-wrong document. Interview angle. “Your support bot gives confidently wrong answers — fix?” A junior reaches for a better encoder or a bigger model; a senior asks whether the candidate set is the problem and adds metadata filtering / source narrowing before retrieval. Reducing the surface area of wrong context is often the highest-leverage safety move in enterprise RAG.

When hybrid isn’t worth it — and how it fails

  1. 01Skip it when queries are pure natural-language paraphrase over a small, clean corpus — you add latency and an index for no gain.
  2. 02Analyzer bug: a BM25 analyzer that lowercases and strips hyphens turns “E-1042” into nothing — your exact-match advantage evaporates. Tune the analyzer.
  3. 03RRF can be dominated by the noisier retriever if it returns far more candidates — cap each list (e.g. 50) and prune.
  4. 04k is a knob: k≈60 is the TREC default; with many relevant docs per query, k=10–30 can help.

The recall stage, end to end

Putting it together, the first stage of a production retriever is a few lines: run both retrievers wide, fuse with RRF, hand the fused shortlist to the reranker (L4). The two details that matter are casting a wide net per retriever (so the gold chunk is present for fusion) and capping each list (so a flood from one retriever can’t skew RRF).
python
1# Hybrid recall -> fuse -> hand off to the reranker (the Perplexity/Uber shape).2def hybrid_recall(query, k_each=50):3    bm25_ids = bm25_search(query, k=k_each)        # exact tokens, codes, names4    dense_ids = vector_search(query, k=k_each)     # paraphrase, intent5    fused = rrf([bm25_ids, dense_ids], k=60)       # rank-based, no calibration6    return fused[:k_each]                          # capped shortlist for the reranker78def retrieve(query):9    shortlist = hybrid_recall(query, k_each=50)    # RECALL: wide, cheap10    return rerank(query, shortlist, top_n=5)       # PRECISION: costly, on shortlist11# Two stages, two jobs: hybrid maximizes the chance the answer is PRESENT;12# the reranker maximizes that the right 5 are at the TOP.

Tuning BM25: k1 and b are not magic constants

If you keep BM25, know its two knobs, because interviewers occasionally probe whether “BM25” means anything more to you than a checkbox. k1 (typ. 1.2–2.0) controls term-frequency saturation — how fast extra occurrences of a term stop adding score; lower k1 saturates sooner (good when repetition is noise, like boilerplate-heavy docs). b (typ. 0.75) controls length normalization — how much long documents are penalized; b near 1 fully normalizes (good for mixed-length corpora), b near 0 ignores length (good when long docs really are more informative). The defaults are sane and rarely the bottleneck — the analyzer is — but knowing what they do separates “I turned on BM25” from “I understand the lexical half of my system.”
articleArchitecting and Evaluating an AI-First Search API (Perplexity’s stack)Perplexity ResearcharticleEnhanced Agentic-RAG: near-human accuracy for on-call supportUber EngineeringdocsReciprocal Rank Fusion — reference & parametersElasticpaperSPLADE: Sparse Lexical and Expansion Model for First-Stage RankingFormal et al. (arXiv)

Checkpoint

A user searches for the exact error code “E1042.” Which retriever most reliably surfaces it?

ADense vector searchBKeyword / BM25, inside a hybrid (with a sane analyzer)CA bigger embedding model
Sign up free to answer and see why

Checkpoint

Why do production systems like Perplexity fuse BM25 + dense and then run a cross-encoder, instead of just doing one bigger vector search?

ABecause vector search is too slow at scaleBRecall and precision are different jobs: fuse cheap recall (BM25+dense), then spend a precise but costly reranker on the shortlistCTo avoid using a vector database at all
Sign up free to answer and see why

Checkpoint

You turned on hybrid, but searches for part numbers like “ACME-X200” still miss even though the docs contain them. Most likely cause?

AThe BM25 analyzer is splitting/stripping the hyphen so “ACME-X200” never becomes a matchable tokenBRRF is weighting dense too heavilyCThe embedding model is too small
Sign up free to answer and see why

Checkpoint

A teammate wants to replace RRF with a tuned weighted fusion (α·bm25 + (1−α)·dense) to squeeze a few NDCG points. What’s the senior caveat?

AWeighted fusion is always worse than RRFBThe weight is corpus/query-distribution-specific, so it goes stale on drift and needs a labelled set plus ongoing re-tuningCWeighted fusion can’t combine BM25 and cosine scales
Sign up free to answer and see why

Checkpoint

Queries are pure natural-language paraphrase over a small, clean FAQ with no codes or proper nouns. A teammate insists on adding BM25 + RRF. Best call?

AAdd it — hybrid is always betterBAdd SPLADE instead, it’s strictly superiorCSkip hybrid here — measure first; pure dense may already be enough for paraphrase-only traffic
Sign up free to answer and see why

Interview prep

Hybrid-retrieval questions test whether you understand that recall and precision are different jobs, that lexical and dense fail on different queries, and that fusion is an operational choice. Interviewers want the mechanism (why BM25 wins exact tokens, why RRF fuses on rank), the production gotcha (the analyzer), and the judgement to know when hybrid isn’t worth it.
  1. 01“Did vector search kill keyword search?” → no — dense compresses away rare exact tokens; BM25 keeps them. They cover different queries, so you fuse.
  2. 02“How does BM25 score?” → IDF-weighted term matches with saturation (k1) and length normalization (b) — not a raw count; still a top BEIR baseline.
  3. 03“Why RRF and not averaging scores?” → BM25 is unbounded, cosine is [-1,1]; fuse on rank so neither scale dominates, and cap each list before fusing.
  4. 04“RRF vs weighted fusion?” → RRF by default (zero maintenance); weighted only with a labelled set and a re-tuning plan, for a few NDCG points.
  5. 05“Why fuse then rerank instead of one big vector search?” → recall (wide cheap net) and precision (costly cross-encoder on the shortlist) are separate stages.
  6. 06“What is SPLADE / learned sparse?” → transformer-predicted sparse weights with term expansion — BM25’s exactness plus some synonym handling, at a larger index.
  7. 07“Hybrid is on but codes don’t match — why?” → analyzer bug (it stripped the punctuation/hyphen); fix tokenization, not fusion.
  8. 08“When do you skip hybrid?” → pure paraphrase over a small clean corpus with no codes/names — measure; the extra index and latency may buy nothing.
Deeper follow-ups separate the strong candidate: “how would you architect hybrid at 200M queries/day?” (one co-located engine like Vespa, not three systems to keep in sync — the Perplexity lesson); “your support bot is confidently wrong” (narrow the candidate set with metadata/source filtering before retrieval, the Uber Genie −60% move, not a bigger model); and “RRF is being dominated by one retriever” (it’s flooding the list — cap candidates per retriever, e.g. 50, and prune). In each, tie the fix to recall vs precision and to maintenance cost.

Could you justify a hybrid + multi-stage design — and when to skip hybrid — in a design review?

FuzzyMostlySolid

Takeaways

  • Dense = meaning/recall; BM25 = exact-term precision (IDF + saturation + length-norm). Fuse with RRF (rank-based, no calibration).
  • RRF beats weighted fusion on maintenance: no corpus-specific weight to go stale on drift.
  • Learned sparse (SPLADE) adds term expansion to BM25 — complement, rarely replacement.
  • Hybrid is the recall stage of a multi-stage ranker — Perplexity/Uber fuse, then rerank, on one co-located engine.
  • Hybrid is near-free after chunking — but the analyzer is the usual failure, and skip it for paraphrase-only corpora.
  • Pruning the document set (Uber Genie, −60% wrong advice) can beat swapping encoders.

Next: the precision stage — cross-encoder reranking, ColBERT, and query transformation.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.