A RAG system has four stages. Ingestion: documents are loaded, chunked (fixed-size, semantic, or hierarchical), embedded with a model like OpenAI text-embedding-3 or BGE-large, and indexed in a vector store (Pinecone/Weaviate/Qdrant/Milvus/Chroma). Query: user query is optionally rewritten (HyDE, multi-query expansion), embedded with the same model, and used to retrieve top-k similar chunks (ANN search via HNSW or IVF); often blended with sparse retrievers (BM25) in hybrid search, then re-ranked with a cross-encoder (Cohere rerank, ColBERT). Augmentation: retrieved chunks are stuffed into the prompt as context, often with instructions to cite. Generation: an LLM produces an answer grounded in the context. Senior additions: (1) chunk metadata filters (date, source, author) can pre-narrow retrieval; (2) parent-doc retriever stores small chunks but returns larger context; (3) late interaction rerankers (ColBERT) win on quality/perf; (4) evaluators run on all four stages independently.