The precision stage: how cross-encoders differ from bi-encoders, reranking (the highest-ROI upgrade), ColBERT late interaction, query transforms, and when GraphRAG, agentic retrieval, and long-context are worth it vs hype.
First-stage retrieval (vector + BM25) is tuned for recall — get the right chunk somewhere in the top 50, fast. But only a few chunks fit in the prompt, and packing it with marginal hits both costs tokens and degrades the generator (more context, more “lost in the middle”). The reranker is the precision stage: a model reads the query and each candidate together and scores true relevance, so you keep the best 3–5.
Cross-encoder reranking: the highest-ROI upgrade in RAG
Why two stages? A bi-encoder (your embedding model) encodes query and docs separately — fast and indexable, but approximate. A cross-encoder encodes them jointly in one forward pass — far more accurate, but ~1000× costlier and impossible to pre-index, so you only run it on the shortlist. The payoff is large and well-measured:
The architectural distinction is the whole interview answer, so make it crisp. A bi-encoder runs the query and each document through the encoder separately and compares the two finished vectors with cosine — so the document vectors can be computed offline and indexed, which is what makes search fast, but the query and document never “see” each other, so the match is approximate. A cross-encoder concatenates query + document into one input and runs full cross-attention over the pair, producing a single relevance score — every query token attends to every document token, which is far more accurate but means you must run a fresh forward pass for every query-document pair and can pre-compute nothing. That’s the ~1000× cost and the reason it only ever runs on a shortlist.
code
1Bi-encoder vs cross-encoder -- why you need both23 Bi-encoder (first stage, recall) Cross-encoder (rerank, precision)4 ------------------------------------ ----------------------------------5 encode(query) -> q vector encode(query + doc) -> 1 score6 encode(doc) -> d vector (OFFLINE) full cross-attention over the pair7 score = cos(q, d) no pre-indexing possible8 ms over millions (ANN) ~ms PER PAIR -> shortlist only9 approximate, scalable accurate, ~1000x costlier1011 Pattern: bi-encoder fetches 50 cheaply; cross-encoder re-scores those 50 to pick 5.
Interview angle. “Why not just rerank everything / why not just use a better embedding model?” is the canonical reranking question, and the answer is this architecture: you cannot pre-index a cross-encoder, so scoring the whole corpus per query is seconds-to-minutes of compute; and a better bi-encoder still can’t match a cross-encoder’s pairwise attention. The two stages exist because recall is cheap-and-approximate and precision is costly-and-exact, and you spend the expensive model only where it matters — on the shortlist.
code
1Adding a cross-encoder reranker over the first-stage shortlist (Cohere Rerank 3.5)23 On a representative financial-services dataset (nDCG):4 Rerank 3.5 vs Hybrid search ..... +23.4%5 Rerank 3.5 vs BM25 .............. +30.8%67 The reranker re-scores the FUSED shortlist for precision, at a few-hundred-ms cost.8 Source: Cohere, "Introducing Rerank 3.5" -- numbers are domain-specific; measure on yours.
python
1import cohere2co = cohere.Client()34candidates = hybrid_search(query, k=50) # wide net (recall)5reranked = co.rerank(6 model="rerank-v3.5",7 query=query,8 documents=[c["text"] for c in candidates],9 top_n=5, # precision: keep the best 510)11context = [candidates[r.index] for r in reranked.results]
Costs and caveats to know: Cohere Rerank 3.5 is ~$2.00 per 1,000 searches at a few-hundred-ms latency; self-hosted ms-marco-MiniLM or bge-reranker run ~50 ms / 100 pairs on a GPU. General rerankers lose 5–10 points on legal/medical/financial — fine-tune on ~10k (query, doc, label) triples before trusting them in a vertical. And rerank the fused list: reranking only the dense top-k throws away BM25’s exact-match hits.
Three production knobs that decide whether the reranker helps or hurts. (1) Shortlist size. Cross-encoder cost is linear in candidates × passage length, so reranking 100 long candidates can add seconds — cap at top-25–50 and truncate passages to ~512 tokens. (2) Cutoff. How many you keep (top-5) trades recall against context budget; too few drops a needed chunk, too many reintroduces noise — tune it on your eval set, don’t guess. (3) Build vs buy. A managed reranker (Cohere, Voyage) is a one-line call but adds a network hop and per-search cost; a self-hosted MiniLM/BGE reranker is ~50 ms/100 pairs on a GPU you operate. Interview angle. “Where does the reranker fit and what does it cost you?” → on the fused shortlist, before prompt assembly; budget ~100–300 ms and either a per-search fee or a GPU, and fine-tune it for any specialized vertical.
Don’t reach for a reranker as the first move, though — the failure it fixes is specific. A reranker only helps when the right chunk is already in your shortlist but ranked too low (failure point #2, “missed top-ranked”). If recall@50 is broken — the gold chunk isn’t retrieved at all — no reranker can recover it, and you’re back to chunking, hybrid, or an ANN-param fix. So the order that works: get chunks reasonable → add contextual + hybrid for recall → then add the reranker for precision. Adding it before recall is solved is a classic wasted optimization.
ColBERT / late interaction — a different first stage
ColBERT encodes every token of the query and document and scores via max-sim (for each query token, the best-matching doc token, summed). It’s ~2 orders of magnitude faster than a cross-encoder because doc embeddings are precomputed; PLAID serves it in tens of ms over 140M passages on a GPU. The cost is storage — a multi-vector index is 10–50× larger (residual compression brings 256 bytes/token down to ~20–36). Crucially, ColBERT replaces your first-stage retriever, not your reranker — reach for it when bi-encoder recall@50 is broken on a 10M+ passage corpus; you can still stack a cross-encoder on top.
The clean mental model: ColBERT is a middle ground between bi-encoder and cross-encoder. A bi-encoder pools everything into one vector (cheapest, least expressive); a cross-encoder runs full cross-attention per pair (most expressive, uncacheable); ColBERT keeps per-token vectors and does a cheap max-sim at query time — more expressive than a single pooled vector, but still pre-indexable. So it captures fine-grained term matches a bi-encoder blurs, without the cross-encoder’s per-pair cost. Interview angle. “Bi-encoder, cross-encoder, ColBERT — when each?” → bi-encoder as the default scalable first stage; cross-encoder as the precision reranker on a shortlist; ColBERT when single-vector recall is genuinely broken at 10M+ passages and you can pay 10–50× index storage. Knowing it replaces the retriever, not the reranker, is the detail that signals depth.
Query transformation: meet the user where they are
The other half of precision is fixing the query, not the index. Real user queries are short, ambiguous, compound, or use words that don’t appear in the corpus — and the embedding of a bad query retrieves bad chunks no matter how good the index is. Query transformation rewrites the query before retrieval. Each technique targets a specific query pathology and each costs at least one extra LLM call, so they’re scalpels, not defaults:
01HyDE — generate a hypothetical answer, embed that, retrieve. Wins on zero-shot / domain-mismatch corpora; HURTS on grounded factual queries (the fake doc drifts from the gold passage).
02Multi-query — fan out 3–5 paraphrases, fuse with RRF. ~5–15% recall on multi-aspect questions; cap fanout at 3 and run it async or latency explodes.
03Decomposition — split a compound question into sub-questions, retrieve each. Needed for “relationship between X and Y in quarter Z” shapes; one LLM call + one retrieval per sub-q.
04Step-back — abstract to a higher-level question first (good for STEM/medicine), then answer with both. Adds an LLM call.
python
1# Multi-query: fan out paraphrases, retrieve each, fuse with RRF (cap fanout!).2def multi_query_retrieve(query, llm, k=20, fanout=3):3 rewrites = llm(4 "Rewrite this query in " + str(fanout) + " different ways, one per line, "5 "covering different phrasings and aspects:\n" + query6 ).splitlines()7 queries = [query] + rewrites[:fanout] # keep the original too8 rankings = [vector_search(q, k=k) for q in queries] # run these ASYNC in practice9 return rrf(rankings)[:k] # fuse -> one ranked list10# Decomposition is the same shape but the LLM splits a COMPOUND question into11# sub-questions, retrieves each, and the answer is composed from all the contexts.
Interview angle. The trap question is “would HyDE help here?” The senior answer is conditional: HyDE shines on zero-shot / domain-mismatch corpora where the query and the documents use different vocabulary, but it hurts on well-grounded factual lookups because the hypothetical answer can hallucinate details that drag retrieval away from the true passage. Naming that it can backfire — and tying each transform to the query pathology it fixes plus its LLM-call cost — beats listing techniques.
Cross-encoder, LLM-reranker, or ColBERT — the reranking menu
There isn’t one kind of reranker, and choosing among them is a real interview question. Cross-encoders (Cohere Rerank, BGE-reranker, MiniLM) are the workhorse: pointwise query-document scoring, cheap enough for ~100 candidates, the default. LLM rerankers ask a general model to score or order candidates — they can be listwise (rank the whole shortlist at once, which captures inter-document relationships a pointwise scorer misses) and they shine on nuanced, instruction-laden relevance (“prefer the most recent, authoritative source”), but they’re slower and pricier, so reserve them for small shortlists and high-stakes queries. ColBERT, as covered above, isn’t really a reranker at all — it’s a first-stage retriever. The senior framing: pointwise cross-encoder by default; listwise LLM reranker when relevance is subjective or instruction-dependent and the shortlist is small; fine-tune any of them for a specialized vertical.
code
1Reranker options -- pick by latency, cost, and how subjective relevance is23 Type Scoring Latency (100) $/search Use when4 --------------- --------- ------------- -------- --------------------------5 MiniLM/BGE (self) pointwise ~50 ms (GPU) your GPU high volume, own infra6 Cohere/Voyage pointwise 100-300 ms ~$2/1k default managed, fast start7 LLM (listwise) listwise seconds LLM tokens subjective/instructional, small list8 Fine-tuned x-enc pointwise ~50-150 ms train once legal/medical/financial verticals910 Default = pointwise cross-encoder. Escalate to listwise LLM only when relevance11 is nuanced AND the shortlist is small enough to afford it.
The reranker also adds a dependency, which is a senior-level consideration juniors skip. A managed reranker is a network call that can time out or rate-limit — so it needs the same reliability discipline as any external call, and a sensible graceful degradation: if the reranker fails, fall back to the fused first-stage order rather than erroring the whole query. Answers get slightly worse, not absent. Interview angle. “What happens when your reranker is down?” → degrade to first-stage ranking and serve a (slightly worse) answer; never let a precision-stage dependency take down retrieval entirely.
python
1# Rerank with graceful degradation: a reranker outage should DEGRADE, not break.2def rerank_safe(query, candidates, top_n=5):3 try:4 scored = cross_encoder_rerank(query, candidates, timeout=0.3) # precision stage5 return scored[:top_n]6 except (TimeoutError, RerankerUnavailable):7 log_metric("rerank_fallback") # alert if this spikes8 return candidates[:top_n] # fall back to fused first-stage order9# Worse ranking beats no answer. Cap candidates (<= ~50) and truncate to ~512 tokens10# so the reranker stays inside its latency budget in the first place.
GraphRAG, agentic & long-context — situational, not default
GraphRAG (Microsoft) extracts an entity-relation graph and pre-summarizes communities, then answers via local search (a subgraph) or global map-reduce over community summaries. It shines for sensemaking — “what are the main themes in this corpus?” — where no single chunk holds the answer. But indexing costs 5–10× a vector index in LLM calls, so it’s wrong for lookups. Microsoft’s dynamic community selection cut global-search cost ~77% (1,500 → 470 communities) at similar quality — but the headline stays: GraphRAG is the global-sensemaking competitor to RAG, not a replacement.
Agentic / iterative retrieval (IRCoT, Self-RAG, FLARE) interleaves retrieval with reasoning. IRCoT adds +21 retrieval / +15 QA points on multi-hop benchmarks — but at 4–7 retrievals per query. For the ~80% of systems doing single-hop factual QA, it multiplies latency for no win. And long-context vs RAG: re-evaluations show long-context generally wins single-hop Wikipedia QA, while RAG wins dialogue, grounded citations, freshness, and cost. The honest answer is “both” — Anthropic’s own guidance is just-in-time retrieval with sub-agent compaction, not dumping the corpus into one context.
The “why not just use a 1M-token context window instead of RAG?” question deserves a sharp answer because it comes up constantly. Long context loses on four axes that matter in production: cost (you pay for every token of the dumped corpus on every query, vs retrieving a few chunks), latency (prefill scales with input length — a full corpus is seconds of TTFT), recall (the Lost-in-the-Middle / context-rot effect means a fact buried in 500k tokens is often not recalled, and gets worse as you fill the window), and attribution (RAG hands you the exact chunk to cite; a giant context can’t tell you which passage it used). Long context is a complement for genuinely holistic, single-document tasks — not a replacement for retrieval over a corpus. Interview angle. Saying “bigger context window solves RAG” is a red flag; the senior framing is “retrieve to keep the window small, cheap, ordered, and citable.”
Why not just run the cross-encoder reranker over the whole corpus and skip first-stage retrieval?
ACross-encoders are less accurate than bi-encodersBIt’s far too slow — a cross-encoder must score every query–doc pair jointly and can’t be pre-indexedCCross-encoders can’t read long text
Users mostly ask narrow factual lookups (“what’s the refund window for plan X?”). A teammate wants to switch the whole system to GraphRAG. Best call?
AYes — GraphRAG is strictly better than vector RAGBNo — keep vector RAG + rerank for lookups; reserve GraphRAG for whole-corpus sensemaking questionsCAdd agentic multi-hop retrieval to every query instead
recall@50 is broken — the gold chunk often isn’t retrieved at all. A teammate adds a cross-encoder reranker to fix it. Will it work?
ANo — a reranker only re-orders the shortlist; it can’t recover a chunk that first-stage retrieval never fetchedBYes — cross-encoders are accurate enough to find anythingCYes, if you also raise top_n to 50
A grounded factual QA system over your own clean docs has good recall. A teammate proposes HyDE to “boost retrieval.” Best call?
AAdd HyDE everywhere — it always improves retrievalBSkip HyDE here — it helps zero-shot/domain-mismatch corpora, but can backfire on well-grounded factual lookupsCReplace retrieval with HyDE-only
An interviewer asks why you don’t just stuff the whole corpus into a 1M-token context and skip retrieval. Strongest answer?
ALong context is cheaper than running a vector DBBIt loses on cost (per-token), latency (prefill), recall (lost-in-the-middle), and attribution (no chunk to cite) — retrieve to keep the window small, cheap, ordered, and citableCModels can’t actually read 1M tokens
Precision-stage questions test whether you know the bi-encoder/cross-encoder/ColBERT trichotomy, whether you add techniques for a reason (the query pathology, the failure point) or by reflex, and whether you can resist the “fancy technique replaces RAG” hype. Interviewers love this lesson because it’s full of “when is X worth it” judgement calls.
01“Bi-encoder vs cross-encoder?” → bi encodes separately (indexable, fast, approximate); cross encodes the pair with full attention (accurate, uncacheable, ~1000× costlier → shortlist only).
02“Why not rerank the whole corpus?” → a cross-encoder can’t be pre-indexed, so it’s seconds-to-minutes per query; first stage narrows millions to ~50 cheaply.
03“What does a reranker buy and cost?” → a large precision lift on the fused shortlist (Cohere reports double-digit nDCG gains over hybrid/BM25) for a few-hundred-ms add; cap candidates at 25–50, truncate to ~512 tokens, fine-tune for verticals.
04“Where does ColBERT fit?” → it replaces the first-stage retriever (10M+ passages, broken single-vector recall) at 10–50× index storage — not the reranker.
05“Reranker not helping — why?” → recall@50 is broken; it only re-orders a shortlist, it can’t recover a chunk that was never retrieved.
06“Would HyDE help?” → only on zero-shot/domain-mismatch corpora; it backfires on grounded factual queries where the fake answer drifts from the gold passage.
07“GraphRAG instead of vector RAG?” → only for whole-corpus sensemaking (“main themes”); on lookups vector RAG + rerank wins on quality and cost, and GraphRAG indexing is 5–10× pricier.
08“Why not a 1M-token context instead of RAG?” → cost, latency (prefill), recall (lost-in-the-middle), and attribution all favour retrieving a small, citable set.
Going deeper, the strong-candidate follow-ups: “order these upgrades for a struggling RAG” (chunks → contextual + hybrid for recall → reranker for precision → query transforms / ColBERT / GraphRAG only if the query distribution demands it — never reranker-first); “your reranker regressed on legal docs” (general rerankers lose 5–10 points on verticals — fine-tune on ~10k labelled triples); “multi-query blew up your latency” (cap fanout at 3 and run paraphrase retrievals async); and “when is agentic/iterative retrieval worth 4–7× the calls?” (genuine multi-hop questions, not the single-hop majority). Match the technique to the failure, and always name its cost.
Could you decide — with reasons — when to add a reranker, ColBERT, query transforms, or GraphRAG?
FuzzyMostlySolid
Takeaways
Bi-encoder (indexable, approximate) vs cross-encoder (pairwise attention, uncacheable, shortlist-only) — that distinction is the reranking interview answer.
Rerank on the fused shortlist: retrieve 50, cross-encode to 5 — a large precision lift (Cohere: double-digit nDCG over hybrid/BM25) for a few-hundred-ms add. Cap candidates, truncate to ~512 tokens.
A reranker only fixes “in the set but ranked low” — it can’t recover broken recall@50. Solve recall first.
ColBERT (per-token, max-sim) is the middle ground — it replaces the first-stage retriever at 10–50× storage, not the reranker.
Query transforms help narrow regimes (HyDE can backfire on grounded queries); each costs an LLM call.
GraphRAG / agentic / long-context are situational — match them to the query distribution, not the hype.
Next: what changes when this pipeline meets billions of vectors, stale indexes, and multi-tenant permissions.