LLM-as-judge uses a (usually stronger) LLM to score, classify, or compare outputs of another LLM. Three modes: (1) pointwise: judge a single answer on a rubric; (2) pairwise: judge "A vs B" which is better; (3) reference-based: judge against a gold answer. Practical recipes: structured rubrics (CoT scoring prompts, criteria-based chains), use a stronger model than the candidate, calibrate against a human-labeled set. Failure modes: (1) position bias: "A vs B" output depends on order; mitigate by averaging both orderings; (2) length bias: longer answers judged better; (3) self-preference: judge favors its own style; (4) format bias; (5) rubric-blind evaluation: LLM judges miss domain-specific errors a human expert catches. Senior view: LLM-as-judge is reliable for ~70-85% of cases but ALWAYS needs human spot-audits for high-stakes claims. Pairwise comparison with win-rate is the most stable metric for online A/B testing.