Lessons
1Why GenAI eval is hard46 min read
The accuracy number you trust for a classifier evaporates here: same input, different output; no single ground truth; and an offline score that lies about production. Why the old metrics break, the eval maturity ladder, and the discipline that replaces them.
- →Evaluation Design
2Golden datasets & rubrics48 min read
The reference data is the eval — and most teams mis-build both halves. Sampling that finds real failures, rubrics that two annotators actually agree on (binary over Likert), the dimensions-and-tuples trick for coverage, and reference-based vs reference-free metrics as a cost-modeling choice.
- →Golden Datasets and Rubrics
- →Evaluation Design
3LLM-as-judge & its traps48 min read
The dominant reference-free metric is itself a biased instrument. Pointwise vs pairwise by criterion shape, the position/verbosity/self-preference biases with measured magnitudes, the prompt mechanics that mitigate them, and why a judge is a probabilistic sensor you must calibrate.
- →Calibrated Model Judges
- →Evaluation Design
4Faithfulness & relevance metrics48 min read
Ragas decomposes RAG failure into orthogonal surfaces — faithfulness, answer relevance, context precision, context recall — each a precise LLM-judge pipeline. Plus the statistic that exposes what percent-agreement hides: Cohen’s kappa, and why it can read 0.3 while Spearman reads 0.9.
- →RAG Quality Measurement
- →Calibrated Model Judges
5Regression suites & monitoring48 min read
Wire evals into CI as deployment gates, monitor four drift vectors in production with confidence intervals, and treat eval cost as an architecture decision — because at scale the eval bill can exceed $1M/month.
- →Regression Gates
- →Production Evaluation
6Capstone: eval pipeline for a chatbot46 min read
Assemble the whole discipline into one trustworthy eval pipeline for a support chatbot — error analysis → golden set → calibrated judge → regression gate → production monitor — and rehearse proving to a skeptical PM that the metric reflects what users actually care about.
- →Evaluation Design
- →Regression Gates
- →Production Evaluation
Skills in this course
- 01Evaluation DesignTurn product quality into explicit, repeatable GenAI evaluation criteria.
- 02Golden Datasets and RubricsBuild representative datasets and reliable annotation rubrics.
- 03Calibrated Model JudgesDesign and calibrate model judges against human labels while controlling known bias.
- 04RAG Quality MeasurementSeparate retrieval, relevance, grounding, and answer-quality failure surfaces.
- 05Regression GatesUse stable eval suites as release gates for model and prompt changes.
- 06Production EvaluationMonitor quality, drift, and eval cost on representative production samples.