Requirements: 1B docs, 100k+ QPS, p95 <100ms, recall@10 >0.9 on labeled queries.
Architecture: ingest (chunk → embed → vector store + BM25 inverted index) → hybrid retriever (weighted sum or reciprocal-rank fusion) → optional reranker → client.
Index internals — know them: HNSW (multi-layer navigable small-world graph; log-ish search, high recall, memory-hungry, awkward to shard) vs IVF-PQ (k-means coarse quantizer + product quantization compressing ~32×; less memory, tunable recall/latency, easier to shard). At 100M+ vectors: shard by IVF cell or random partition; fan out, gather, merge.
Freshness: ANN indexes hate in-place updates — run a small "fresh" brute-force/HNSW buffer for recent items alongside the big immutable index, merge results, rebuild on a schedule.
Eval: nDCG@10 against human-labeled queries; offline recall@k vs latency curve; recall@k vs exact brute force on a sample.
Monitoring: index health (orphan vectors, stale fields), recall trend by query class, drift in embedding distribution (cosine to an anchor set).
Cost/latency: quantize embeddings (int8) for memory; cache popular queries; offload reranking to a smaller model.
Follow-up probes: How do you re-index without downtime? (Dual-index blue/green, backfill, evaluate before cutover.) Multi-lingual? What happens when the embedder changes? Why not brute force on GPUs? (Legit to ~10M vectors — saying so shows judgment.)