RAGAS (Reference-free Augmented Generation Assessment) is an open-source framework for RAG eval that uses LLM-as-judge in four canonical metrics: faithfulness (does the answer stay in retrieved context?), answer_relevancy (does it answer the question?), context_precision (are the retrieved chunks actually relevant?), and context_recall (does the corpus contain answers to the question?). It became popular because it works without ground-truth labels: the LLM judge derives scores from question + context + answer alone. Limits: quality depends on the LLM judge, so biases in the judge propagate; not great for very long-form or numeric answers; some metrics assume a "single right answer" when queries are open-ended. Senior nuance: pair RAGAS with a small labeled set to ground the LLM judge and detect drift. Newer tools (DeepEval, Arize Phoenix, TruLens) extend RAGAS with traces, hallucination detection, and pairwise tests.