Ragas decomposes RAG failure into orthogonal surfaces — faithfulness, answer relevance, context precision, context recall — each a precise LLM-judge pipeline. Plus the statistic that exposes what percent-agreement hides: Cohen’s kappa, and why it can read 0.3 while Spearman reads 0.9.
Decompose the failure, then trust the agreement
A RAG system has three independent failure surfaces — the retriever (did we fetch the right context?), the reader (did the model reason correctly over it?), and the answer surface (is this what the user asked?). One end-to-end accuracy number conflates all three and tells you nothing about where it broke. Ragas gives you a metric per surface, each a precise LLM-judge pipeline. And because every one of those metrics is judge-driven, this lesson closes with the statistic that tells you whether to believe a judge at all: Cohen’s kappa.
The senior framing, and a classic interview trap: faithfulness and answer relevance are orthogonal. Faithfulness asks "does the answer say only what the retrieved context supports?" — a hallucination guard. Answer relevance asks "does the answer address the user’s intent, regardless of source?" — a usefulness guard. They factor into a 2×2: faithful + relevant is the ideal; not-faithful + relevant is the highest-risk mode (confidently wrong, on-topic, plausible); faithful + not-relevant means the model echoed retrieved text without answering; neither is off-topic and unsupported. Candidates who collapse the two into "is the answer good?" lose this round — the whole point is that the fixes differ.
Faithfulness: did the reader hallucinate against context?
Faithfulness measures how factually consistent a response is with the retrieved context, scored 0–1, higher better. The Ragas pipeline is mechanical and worth knowing exactly: (1) break the generated answer into individual claims; (2) for each claim, the LLM judge verifies whether it can be inferred from the retrieved context; (3) score = (claims supported by context) / (total claims). It needs user_input, response, and retrieved_contexts — and is reference-free for the answer (no gold answer required). The critical property: a faithfulness of 1.0 does not mean the answer is true — it means the answer is grounded in the retrieved context, no more. A confidently-wrong-but-cited answer passes faithfulness; what it catches is unsupported claims, the hallucination surface.
python
1# Faithfulness, by hand: decompose into claims, check each against context.2def faithfulness(answer, contexts, judge):3 claims = judge(f"List the atomic factual claims in:\n{answer}\nOne per line.")4 claims = [c for c in claims.splitlines() if c.strip()]5 supported = 06 for c in claims:7 verdict = judge(8 "Can this claim be inferred from the context? Answer YES or NO.\n"9 f"{chr(10).join(contexts)}\n{c}"10 )11 supported += 1 if verdict.strip().upper().startswith("YES") else 012 return supported / max(len(claims), 1)13# 1.0 means "grounded in context", NOT "true". Catches unsupported claims only.14# The claim-check step is itself a judge -> it needs its own TPR/TNR calibration.
Answer relevance: did it answer the question asked?
Response Relevancy measures how pertinent the answer is to the prompt, and its pipeline is the most distinctive of the four: (1) generate N artificial questions whose answer could plausibly be the model’s response (counterfactual generation, a judge step); (2) embed each generated question and the original user input; (3) score = mean cosine similarity between the generated questions and the original. The intuition: a highly relevant answer is one that, reverse-engineered into questions, yields questions that look like the one actually asked. Ragas’s own example — Q: "Where is France and what is its capital?"; low-relevance answer "France is in western Europe." (partial); high-relevance "France is in western Europe and Paris is its capital." It catches the "LLM answered a different question" mode common with ambiguous retrieval, independent of factual correctness — a perfectly factual but unrelated reply scores low.
Context precision & recall: the retriever’s two failure modes
Context Precision asks whether the relevant chunks are ranked above the irrelevant ones: for each chunk at rank k the judge scores it relevant/not against the question and ground-truth answer, and the metric accumulates a rank-weighted precision. Its use case is retriever regressions — if a chunk-splitter change drops a key clause to rank 8 in the top-10, context precision moves before downstream answer quality does, making it a leading indicator of faithfulness regressions. Context Recall asks whether the retrieved context covers the ground truth: the reference answer is decomposed into claims, and recall = fraction of those claims attributable to a retrieved chunk. It needs the reference (gold answer) and catches missing-context failures — the inverse of precision.
code
1Ragas: four metrics, four orthogonal failure surfaces23 Metric Catches Inputs Ref?4 ---------------- ---------------------- ------------------------- ----5 Faithfulness hallucination vs context user_input, response, ctx no6 Answer Relevance off-topic / wrong-Q user_input, response no7 Context Precision ranker regression user_input, ctx, reference yes8 Context Recall missing-context / gap user_input, ctx, reference yes910 Deploy all four: a Context Recall drop with stable Faithfulness = retrieval11 miss (tune the retriever); a Faithfulness drop with stable Recall = reader12 hallucination (force citations / shrink context). Different levers.
The catch that the README hides: every Ragas metric is an LLM-judge pipeline, so each inherits judge error and needs its own calibration. The claim-extraction and claim-verification steps in faithfulness, the question-generation in relevance, the per-rank relevance scoring in precision — all are judge calls with their own TPR/TNR. A small or weak judge model gives a low true-positive rate at the inference step and silently deflates the metric. Interview angle. "Your Ragas faithfulness reads 0.92 — do you trust it?" → only after validating the underlying judge against a human-labeled sample; the framework score is only as good as the judge under the hood, and Ragas does not calibrate that judge for you.
Cohen’s kappa: why percent agreement lies
Now the statistic that governs whether any of these judges can be trusted. When you say "the judge agrees with humans 85% of the time," you’ve stated something partially correct and mostly misleading. Percent agreement is the fraction of cases two raters label identically — easy to compute, easy to misread, because with class imbalance it can exceed 90% while the judge systematically misses the minority (failure) class. The "Judging the Judges" paper documents a 30-point spread between judges under percent agreement that compresses to chance. Cohen’s kappa fixes this by discounting the agreement expected from random labeling: kappa = (P_o − P_e) / (1 − P_e). The same predictor set that showed a 30-point spread under percent agreement shows a 53-point spread under kappa, because kappa strips out the agreement everyone gets for free by picking the majority class.
code
1Cohen's kappa: the agreement number that doesn't lie23 Interpretation (Landis-Koch) Eugene Yan's operating bands4 -------------------------------- ----------------------------5 < 0.0 worse than chance 0.4 - 0.6 substantial6 0.0-0.2 slight > 0.7 excellent7 0.2-0.4 fair8 0.4-0.6 moderate / substantial TriviaQA-style (96% human IAA):9 0.6-0.8 substantial GPT-4 Turbo k = 0.8410 0.8-1.0 almost perfect Llama-3 70B k = 0.7911 Llama-3 8B k = 0.62 <- substantial, NOT12 kappa = (P_o - P_e) / (1 - P_e) almost-perfect: shipping a13 Report BOTH: percent for headline, judge here understates the14 kappa for the operating point. noise floor by an order of mag.
Pick the statistic to the type of label the judge produces. Binary pass/fail → Cohen’s kappa. Ordinal ratings (1–5) → quadratic-weighted kappa or Spearman’s rho. Multiple annotators (>2) with mixed scales → Krippendorff’s alpha. And know the trap that confuses teams: kappa and rank-correlation can disagree mechanically. Eugene Yan reports a case where Cohen’s kappa was 0.3–0.5 (fair) while Spearman’s rho and Kendall’s tau read 0.8–0.9 — not a contradiction: on a small, imbalanced label set, kappa punishes minority-class errors harshly while rank-ordering rewards consistent ordering. A high-rho / low-kappa pattern is a warning that the judge is mostly defaulting to the majority class.
Interview angle. The single highest-signal exchange in an eval interview: "We have 90% agreement with our judge." → "90% percent agreement, but what’s the Cohen’s kappa? On imbalanced labels 90% can mean kappa near zero — the judge is barely beating a coin flip on the class you care about, so I’d re-stratify to balance the labels and re-measure." Saying that — report two numbers, percent for the headline and kappa for the operating point, never percent alone — puts you in the top quintile of candidates. The same calibration the previous lesson framed as TPR/TNR is here expressed as kappa; both are ways of refusing to trust a raw agreement rate.
Answer correctness: when you do hold a gold answer
The four metrics so far are reference-free for the response (faithfulness, relevance) or reference-light on the retriever. But when you do have a gold answer, you can grade the answer surface directly with answer correctness — Ragas combines a factual-similarity component (how well the answer’s claims match the gold answer’s claims, an F1 over claim overlap) with a semantic-similarity component (embedding similarity to the gold). This is the reference-based generation-side metric, and it answers a question faithfulness can’t: faithfulness says "grounded in retrieved context," correctness says "matches the truth." The research is sharp on when to prefer it: reference-guided judges like Prometheus degrade most when the reference answer is removed, so reference-based grading is the better choice wherever gold coverage exists, and reference-free is the production fallback where it’s thin.
code
1Generation-side metrics: pick by whether you hold a gold answer23 Metric Needs gold? Answers the question Use when4 ---------------- ----------- ------------------------- --------------------5 Faithfulness no grounded in context? production, no gold6 Answer relevance no addressed the question? production, no gold7 Answer correctness yes matches the truth? you have gold answers8 = factual similarity (claim F1) + semantic similarity (embedding)910 Reference-based (correctness) is sharper but bounded by gold coverage;11 reference-free (faithfulness/relevance) generalises but inherits judge bias.12 Senior move: correctness on the stable golden set, faithfulness on live traffic.
Your RAG bot returns answers that are "fluent but sometimes wrong on numbers." Faithfulness is 0.95 and answer relevance is high. Where do you look next, and why?
AContext recall — high faithfulness means the model stuck to context, so wrong numbers imply the right figures weren’t retrievedBFaithfulness is the problem — lower the threshold to catch more hallucinationsCAnswer relevance — make the answers more on-topic
A teammate reports "our LLM judge has 91% agreement with human labels, so it’s reliable." The label set is 94% pass / 6% fail. What’s the issue?
ANo issue — 91% agreement is high, ship itBPercent agreement is inflated by the imbalance; report Cohen’s kappa — it may be near zero, meaning the judge barely beats chance on the rare failure class. Re-stratify and re-measureCSwitch to Spearman correlation, which is always more reliable than kappa
You report Ragas faithfulness = 0.92 to leadership as proof the system rarely hallucinates. What caveat must accompany it?
ANone — 0.92 faithfulness directly means a 92% truthfulness rateBFaithfulness means grounding in context (not truth), and the score comes from an LLM-judge pipeline whose own TPR/TNR must be validated against human labels firstCJust note it’s an average — report the median instead and it’s fine
After a chunk-splitter change, end-to-end answer quality looks unchanged but you want an early warning of retrieval regressions. Which metric is the leading indicator?
AAnswer relevance — it moves first when retrieval degradesBFaithfulness — it catches the regression before anything elseCContext precision — it detects a relevant chunk dropping in rank before downstream answer quality moves
You’re validating a binary pass/fail judge and a colleague computes Spearman’s rho = 0.88 as the agreement metric. What’s the correct critique?
AFor binary labels use Cohen’s kappa (chance-corrected); Spearman is for ordinal/ranked labels, and a high-rho/low-kappa split would itself flag majority-class defaultingBSpearman 0.88 is excellent — nothing to critiqueCUse raw percent agreement instead — it’s simpler and sufficient
Faithfulness/agreement rounds test whether you decompose RAG failure into orthogonal surfaces and whether you validate judges with the right statistic. Interviewers probe four things: do you treat faithfulness and relevance as orthogonal, can you describe each Ragas pipeline mechanically, do you know every Ragas metric is an uncalibrated judge, and do you reach for Cohen’s kappa over percent agreement. Use numbers.
01“Faithfulness vs relevance?” → orthogonal: faithfulness = grounded in context (hallucination guard), relevance = addresses the question (usefulness guard); not-faithful + relevant is highest risk.
02“How does faithfulness work?” → decompose answer into claims, judge each against context, score = supported/total; reference-free, 1.0 ≠ true.
03“How does answer relevance work?” → generate counterfactual questions from the answer, embed, cosine-sim to the original input; catches "answered a different question."
04“Context precision vs recall?” → precision = relevant chunks ranked high (leading indicator of retrieval regressions); recall = retrieved context covers the gold claims (missing-context).
05“Do you trust Ragas faithfulness 0.92?” → only after validating the underlying judge’s TPR/TNR; Ragas inherits judge bias and doesn’t calibrate it for you.
06“Why not percent agreement?” → imbalance inflates it (90% with kappa~0); report Cohen’s kappa for the operating point, percent only for the headline.
08“kappa 0.3 but Spearman 0.9 — what gives?” → small imbalanced set: kappa punishes minority-class errors, rank-order rewards consistency; the judge is defaulting to the majority class.
Going deeper, the follow-ups: "your faithfulness judge is monotone in length — how do you mitigate?" (instruct it to locate claims first then check each, clip over-long answers, calibrate expected length per query type); "95% faithfulness but 60% answer relevance — where first?" (retrieval recall@k stratified by query type — high faithfulness means the model stuck to context, so the missing 35% is context not containing what was asked); and "only budget for one judge in production — which?" (faithfulness, because unsafe/reputational failures are usually faithfulness failures and relevance degrades gracefully into CSAT/re-open). Tie each to the metric and the lever.
Could you localise a RAG failure with the four Ragas metrics and validate the judges with the right agreement statistic?
New to itGetting thereConfident
Takeaways
Faithfulness and answer relevance are orthogonal — grounding vs usefulness; not-faithful + relevant is the highest-risk mode.