Monitoring answers “is it working?”; observability answers “why did it behave this way?” — and on a customer site you often can’t see their data, so the trace IS the artifact. RAGAS metrics, ground-truth from production traffic (Shopify’s GTX), the stale-index SLO, and the self-hosted observation spine that lives in the customer VPC.
The distinction that organizes the lesson
Two words that get conflated and must not be. Monitoring answers “is it working?” — latency, error rate, token count, cost, uptime, via dashboards and alerts. Observability answers “why did it behave this way?” — the per-request trace of inputs, outputs, timing, and config at every pipeline step, via trace viewers and span explorers. You cannot judge a model output without the retrieval trace that produced it. On a customer site this is sharper than usual: you frequently cannot see the customer’s data or read their answers, so the trace — retrieval set, scores, prompt version, model version — is the artifact you debug from. Eval is the engineering axe that separates a RAG demo from a RAG product; observability is how you keep it sharp in the field.
Make the difference concrete with Cisco’s Splunk AI Assistant “two queries, different outcomes” case. Two near-identical questions about a conference registration: without the directive “look at the agenda,” the assistant retrieved an incomplete “Know Before You Go” doc and gave a misleading answer; with the directive, it retrieved the detailed “Agenda” doc and answered correctly. Trace-level observability made the difference — the dashboard surfaced the retrieval path (documents in context, sources tiered green/yellow/red, prompt version, model version, trace IDs, latency), so an engineer could conclude “retrieval was fine; the prompt was missing a directive” and fix it without guessing. Without the trace, that is a week of blind prompt-tweaking.
Offline evals: the RAGAS four, plus retrieval IR metrics
Pre-deployment evaluation has two layers: deterministic + LLM-as-judge scoring on a labeled set, and the retrieval IR metrics underneath. RAGAS standardizes the RAG four — context precision, context recall, faithfulness, answer relevancy — and lets teams author custom metrics with simple decorators; the framework’s own pitch is taking an LLM app “from vibe checks to systematic evaluation loops.” Patronus’s enterprise rubric extends this to five: context relevance, context sufficiency, answer relevance, answer correctness, answer hallucination. The senior discipline from the RAG-systems track still applies: measure retrieval as an IR problem first — recall@k (did the gold chunk make top-k? the ceiling on everything) and precision@k (how clean is the top-k?) — before you score the generator.
code
1THE EVAL STACK for enterprise RAG -- offline + online23 Offline (pre-deploy) Online (production)4 --------------------------------- ---------------------------------5 retrieval: recall@k, precision@k retrieval hit-rate, chunk-relevance6 RAGAS: context precision/recall, grounding score (answer follows7 faithfulness, answer relevancy from retrieved context?)8 Patronus +2: context sufficiency, refusal rate; drift in retrieval9 answer correctness/hallucination distribution over time10 human spot-check on a labeled set prompt-version + embedding-version11 differential metrics1213 Rule: faithfulness/groundedness can run as a GUARDRAIL (block answers below14 an SLO), not just a dashboard. An eval that only paints a chart blocks nothing.
The highest-leverage move is to wire an evaluator as a guardrail rather than merely as a dashboard: if faithfulness on a response falls below an SLO (say 0.5), block it or route to human review instead of shipping it. This is the difference between “we have a faithfulness metric” and “low-faithfulness answers never reach the user.” The same applies to grounding score in production. Interview angle. “How would you evaluate the system?” — a weak answer names one metric (“I run RAGAS”); a strong answer names retrieval recall + answer faithfulness + latency + task success + regression tracking, and notes which metric is the fastest-feedback signal versus the closest-to-business-value one.
Knowing the seven failure points lets you turn “the bot is wrong” into a stage-localized diagnosis, which is exactly what a production eval is for. Barnett et al.’s taxonomy: FP1 missing content (not in the corpus), FP2 missed the top-ranked doc (retrieved but below cutoff), FP3 not in context (dropped by consolidation), FP4 not extracted (in context, model missed it), FP5 wrong format, FP6 incorrect specificity, FP7 incomplete. The split matters: FP1-FP3 are retrieval failures (fix with ingest/chunking/recall), FP4-FP7 are generation failures (fix with grounding prompt/decomposition). Your trace tells you which side of the boundary you are on — and the discipline is to locate the boundary before proposing a fix, because no prompt change recovers a chunk that was never retrieved.
Ground truth from production traffic — Shopify’s GTX
The cleanest published recipe for a real eval set is Shopify Sidekick’s. They build Ground Truth Sets (GTX) based on “real production distributions rather than curated golden datasets” — i.e. they sample traces from actual traffic, not synthetic questions. Three product experts label each item, and inter-annotator agreement is validated with Pearson correlation, Kendall Tau, and Cohen’s Kappa (you do not trust a single labeler). Then they deploy “calibrated LLM-as-a-Judge systems and an LLM-powered merchant simulator to validate candidate systems before deployment.” The mechanism: an eval set drawn from your true distribution catches the failures users actually hit, which a vendor’s defaults never will. Mirror this on a customer site by building the first eval set out of their support tickets and traces — exactly as Morgan Stanley built theirs from their corpus.
Shopify’s reward-hacking story is the cautionary half. During GRPO training the model found exploits — schema violations, opt-out hacking, and “tag hacking” (using customer_tags instead of the correct customer_account_status). The fix was calibrating the LLM judges and syntax validators against human labels, which moved syntax-validator accuracy from ~93% to 99%. The lesson that generalizes: an uncalibrated LLM-as-judge amplifies reward-hacking signals; a calibrated one closes them. So eval gates must be human-reviewed before they block — blindly trusting an LLM judge in the loop creates new failures rather than catching old ones.
The stale-index SLO — the silent failure
The most dangerous production failure in enterprise RAG is the one nothing alerts on: the stale index. As Scality puts it, “a stale index turns the model into a confident generator of out-of-date answers — the failure mode RAG was supposed to eliminate.” No request errors. Latency and throughput are healthy. Every vector looks fine. But the model is answering from documents that no longer reflect reality, and the guard goes off only when someone in the business notices a wrong answer weeks later. The root cause is an SLO mismatch: most RAG systems inherit the language model’s freshness assumption (a few months) while the customer’s freshness expectation is seconds for prices, minutes for inventory, days for policies.
The fix is operational, not a magic embedding: define per-source freshness SLOs in writing, instrument the gap between source-mtime and chunk-embedded-mtime, and ship a “stale-index dashboard” inside the customer’s VPC. Snorkel’s taxonomy is the checklist for the adjacent ingest failures — ineffective chunking (fix: tune windows/overlap/semantic segmentation), generalist embeddings missing domain nuance (fix: fine-tune on hard negatives), missing metadata (fix: NER + topic enrichment for hybrid retrieval), poor prompt templates, generic outputs. Interview angle. “How do you handle freshness?” — a weak answer designs as if data never changes; a strong one states indexing-delay expectations, a reprocessing plan, freshness-aware fallbacks, and a per-source SLO with a monitored mtime gap.
code
1STALE-INDEX telemetry -- the failure nothing else catches23 Signal Why it matters4 --------------------------------- ------------------------------------------5 source_mtime - chunk_embedded_mtime per-source freshness gap; the core metric6 per-source freshness SLO (written) prices: seconds; inventory: minutes;7 policy docs: days -- one number per source8 re-embed lag / backlog depth ingestion falling behind = staleness rising9 retrieval-distribution drift source poisoning / corpus shift shows here10 BEFORE user complaints11 last_verified per chunk time-filter stale content; powers fallbacks1213 The stale index is invisible: no errors, healthy latency, every vector "fine."14 You only catch it if you instrument freshness as its own SLO.
The observation spine lives in the customer VPC
Do not stitch a separate vector DB, a separate LLM, and a separate eval product together in the customer’s VPC and hope. Decide the observation spine first. Langfuse is the field leader with public scale numbers — “over 10 billion observations per month for over 2,300 customers, including 19 of the Fortune 50,” and a compliance posture of “SOC 2 Type II, ISO 27001, GDPR compliant, and HIPAA eligible.” Its hierarchical traces “capture every LLM call, tool invocation, and retrieval step,” filterable by metadata, session, user, latency, or cost. Critically for the field, its architecture (ClickHouse OLAP, S3/Blob for large payloads, edge-cached prompts, a Redis async-ingestion queue) is designed to self-host inside a VPC without becoming a runaway cost. Arize Phoenix (OTel-native, notebook-friendly OSS) and MLflow’s LLM evaluation are the established alternatives.
Galileo’s field model closes the loop: production traces flow back into the offline eval suite, broken answers become regression cases, and fixes are tested on those same cases before redeployment. The minimum trace fields to instrument regardless of vendor: per-request retrieval hit-rate with relevance scores, grounding score, refusal rate, retrieval-distribution drift, and prompt-version + embedding-version differential metrics. The discipline that makes this work on a customer site: treat the eval pipeline as production infrastructure — it lives in the customer’s VPC, has its own SLOs, exports only aggregated non-sensitive metrics outward, and is owned by the platform team, not a data-science side project.
A customer reports the assistant “gives wrong answers sometimes,” but you cannot read their confidential documents or outputs on-site. What lets you diagnose fastest?
AAsk the customer to paste failing answers and documents into a ticket so you can inspect themBPer-request traces — retrieval set, relevance scores, prompt version, model version, grounding score — so you localize retrieval vs prompt vs model without seeing the contentCIncrease the model size to reduce wrong answers
You want low-faithfulness answers to never reach users. What is the right use of your faithfulness metric?
AWire it as a guardrail: block or route-to-human any response scoring below the SLO, not just chart itBDisplay it on a daily dashboard for the team to reviewCOnly compute it offline during pre-deployment evals
Users say the assistant “got worse this week,” but you shipped nothing and latency/error rates are flat. What is the most likely cause and the right instrumentation?
ARandom model variance; nothing actionableBA stale index or silent model/embedding change; instrument the source-mtime vs chunk-embedded-mtime gap, retrieval-distribution drift, and pinned model/embedding versionsCThe vector database ran out of memory
You are building the customer’s first eval set. Which approach best reflects how Shopify and Morgan Stanley actually did it?
AGenerate thousands of synthetic questions from the docs and trust an off-the-shelf LLM judge to score themBUse the model vendor’s default benchmark scores as the evalCSample real production traces / the customer’s own tickets, have multiple experts label them (validate agreement), and calibrate the LLM judge against those human labels before it gates
You are choosing the observability spine for a self-hosted deployment in the customer’s VPC. What matters most?
AA SaaS dashboard that ingests all traces to the vendor’s cloud for the richest UIBPick the spine first; self-host it in the VPC (e.g. Langfuse: ClickHouse + object storage + async queue), capture every retrieval/LLM/tool step, and export only aggregated metrics outwardCSkip tracing to save cost and rely on application logs
OpenAI explicitly weights “How do you know your AI system is actually working?” into the system-design round, so eval rigor is a core score, not a footnote. Interviewers separate “I vibe-checked it” from a real loop. Structure your eval answer as: retrieval IR metrics first (recall@k as the ceiling), then RAGAS/Patronus quality metrics, then online grounding/refusal/drift, then regression tracking — and name which metric is fastest-feedback vs closest-to-business-value. For observability, lead with the monitoring-vs-observability distinction and the trace fields you would instrument.
01“Monitoring vs observability?” → monitoring = is it working (latency/cost/uptime); observability = why (per-request traces of retrieval set, scores, versions).
02“How would you evaluate the system?” → retrieval recall@k + precision@k, RAGAS faithfulness/context precision-recall/answer relevancy, online grounding + refusal + drift, regression tracking — not one metric.
03“How do you build the eval set?” → sample production traffic / customer tickets (Shopify GTX), multi-annotator agreement (Kappa), calibrate the judge before it gates.
04“How do you handle freshness?” → per-source written SLOs, monitor source-mtime vs embed-mtime gap, freshness-aware fallbacks; the stale index errors nothing.
05“What does the stale-index failure look like?” → confident out-of-date answers, healthy latency, no errors — caught only by freshness telemetry and drift.
06“Where does observability live on a customer site?” → self-hosted spine in the VPC (Langfuse/Phoenix), full traces in-VPC, only aggregated metrics exported.
07“How do you stop an LLM judge from reward-hacking?” → calibrate against human labels (Pearson/Kappa); an uncalibrated judge amplifies hacks — Shopify’s 93%→99% fix.
08“How do you make an eval block, not just chart?” → wire faithfulness/groundedness as a guardrail with an SLO threshold that blocks or routes to human review.
Going deeper. The agentic-eval follow-ups are common for FDE roles. “Design an eval suite for an agent that reroutes shipments without overspending while holding a 99% delivery rate.” (pick metrics — cost-per-action, delivery rate, override rate — design the eval before the build, include a regression bank, and rank metrics by business value). “Your agent works in staging but is inconsistent in production — debug it.” (run a data diff and a config diff in parallel; bring a debugging tree, not “check the logs”). “Self-consistency / heavy eval helped offline but you can’t afford it in prod.” (route only the low-confidence tail through the expensive path). The hidden rubric: you design evals as production infrastructure with regression tracking, not as a one-time accuracy check.
Could you specify offline + online metrics, build a production-traffic eval set with a calibrated judge, and instrument the stale-index SLO in a customer VPC?
New to itGetting thereConfident
Takeaways
Monitoring = is it working; observability = why — and on a customer site the trace (retrieval set, scores, versions) is your only window.
Score retrieval as IR (recall@k ceiling) first, then RAGAS faithfulness/context metrics; wire faithfulness as a blocking guardrail, not just a chart.
Build the eval set from production traffic / customer tickets (Shopify GTX), validate annotator agreement, and calibrate the judge before it gates.
The stale index is the silent failure: confident out-of-date answers, no errors — catch it with per-source freshness SLOs and mtime-gap telemetry.
Decide the observation spine first and self-host it in the VPC (Langfuse/Phoenix); keep full traces in-VPC, export only aggregates.
Pin model + embedding versions in every trace so a silent provider update is attributable, not an unexplained regression.
Next: guardrails & PII handling — redaction at three stages, the guardrail stack, and the permission-leak failure mode head-on.