Requirements: 1M docs; private corpus (PII/redaction concerns); sub-3s p95 latency; refusal when context is empty is acceptable.
Architecture: ingestion (loader → parser → chunker — recursive/structural → embedder → vector index) + serving (query embed → hybrid retriever top-50 → reranker top-5 → prompt assembler → LLM → streamed answer with citations).
Data: chunk overlap 100–200 tokens; preserve section/heading metadata; embed with a domain-tuned model or fine-tune on contrastive pairs from the corpus.
Index: Qdrant/FAISS/pgvector + BM25 with reciprocal-rank fusion; rerank via cross-encoder or LLM yes/no. Apply ACL filters at query time in the index, never post-LLM.
Eval: offline — nDCG@10 against human-labeled Q-A pairs, retrieval recall@k, generation groundedness via NLI or LLM judge, answer-F1 against labeled spans; online — thumbs, follow-up rate. Measure retrieval and generation separately: they need different fixes, and conflating them is the classic junior mistake.
Monitoring: retrieval recall over time, drift in chunk-length distribution, judge-vs-human agreement, latency p50/p95 by component.
Cost/latency: prefix caching across system prompts; reranker only on the head, not the tail; dynamic embedding batches; on empty context — refuse rather than hallucinate.
Depth signals: reranking lifts nDCG by 5–15 points and is the "second-best lever" after chunking; refusal + citations are the safety cornerstone.
Follow-up probes: How do you keep the index fresh when docs change hourly? What happens when the embedder is replaced? Multi-tenant isolation? How do you evaluate groundedness reliably?