Lessons
1Retrieval foundations & the shapes of RAG46 min read
Why retrieval — not the model — caps your whole system, how embeddings and ANN actually work, the embedding choices that move the needle, and the three production shapes of RAG you must pick between before you write a line.
- →Embeddings and Vector Search
- →Production RAG Architecture
2Chunking & contextual retrieval48 min read
The unglamorous decision that silently caps every RAG system — chunking failure modes, sizing by metric, parsing as the real bottleneck, late chunking, Anthropic’s contextual retrieval, small-to-big, and how systems like Cursor chunk at scale.
- →Chunking and Context
- →Production RAG Architecture
3Hybrid retrieval & the multi-stage pipeline44 min read
Dense misses exact tokens; BM25 misses paraphrase. How BM25 actually scores, why you fuse with RRF, the multi-stage pipeline Perplexity and Uber run in production, sparse-neural retrievers, and when hybrid isn’t worth it.
- →Hybrid Retrieval and Reranking
4Reranking, ColBERT & query transformation50 min read
The precision stage: how cross-encoders differ from bi-encoders, reranking (the highest-ROI upgrade), ColBERT late interaction, query transforms, and when GraphRAG, agentic retrieval, and long-context are worth it vs hype.
- →Hybrid Retrieval and Reranking
- →Production RAG Architecture
5RAG at scale: index economics, freshness & multi-tenancy50 min read
What breaks only at scale — HNSW/IVF/IVFPQ memory economics, the build-vs-buy decision, the silent freshness graveyard, multi-tenant permission leaks, and caching/cost — with case studies from Perplexity, Harvey, Notion and LinkedIn.
- →Production RAG Architecture
- →Embeddings and Vector Search
6Evaluating & operating RAG48 min read
Prove it works without fooling yourself: the IR retrieval metrics, split retrieval vs generation eval, where RAGAS metrics lie, LLM-as-judge biases & calibration, golden sets, observability, drift, and prompt-injection in the retrieval path.
- →RAG Evaluation
- →Production RAG Architecture
7Capstone: design a production RAG44 min read
Assemble the whole pipeline into a permission-aware, cited, evaluated assistant over private docs — the full ingest-to-answer architecture, the numbers to defend, and a rehearsal of the decisions an interviewer (or an incident) will push on.
- →Production RAG Architecture
- →RAG Evaluation
Skills in this course
- 01Production RAG ArchitectureDesign a grounded, permission-aware retrieval pipeline from ingestion through answer generation.
- 02Embeddings and Vector SearchSelect embeddings and vector indexes using measured recall, memory, and freshness needs.
- 03Chunking and ContextParse and chunk source material so retrieved evidence stays complete and identifiable.
- 04Hybrid Retrieval and RerankingCombine lexical, dense, and reranking stages to improve recall and precision.
- 05RAG EvaluationMeasure retrieval and generation separately and use reliable release gates.