Requirements: support multi-model A/B, multi-judge (human + LM + heuristic), 10K-prompt test suites at <10% of model-serving cost, reproducibility under model change.
Architecture: dataset registry (huggingface-hub style), task runner (Ray/K8s), judge service returning (score, rationale), results store (ClickHouse), dashboards (nDCG, accuracy, calibration, drift).
Judges, in order of trust: exact/programmatic checks where possible (format, citations present, code runs) → model-based rubric grading with chain-of-thought → periodic human audit of judge agreement. LLM-as-judge hygiene: pairwise beats absolute scores; randomize position (position bias is real); use a judge model different from the generator (self-preference bias); calibrate against human labels and report judge-human agreement, not just judge scores.
CI gate: every prompt/model change runs the suite; block on regression beyond a threshold with statistical significance (bootstrap CIs — 3 points on 200 cases can be noise); track per-slice results so an aggregate win can't hide a slice regression. Example ship rule: "release iff judge score does not regress by >0.5% AND human spot-check matches 95%."
Monitoring: per-prompt drift, per-task score drift, calibration drift.
Cost: batch judge calls; cache by (prompt_hash, model_id); tier "release gate" (fast) vs "audit" (slow) suites.
Follow-up probes: How do you avoid the judge being gamed by the model it scores? How do you detect data leakage between training and eval? How do you handle multi-criteria trade-offs (latency vs quality)? Evals for agents/multi-turn? (Trajectory-level: task success, tool-call correctness, steps-to-completion.)