Lesson 7 of 8 · 52 min

Logs, metrics, traces, and on-call for backends

RED/USE, SLIs/SLOs and error budgets, structured logs, cardinality, tracing across queues, alerting, and LLM-specific ops signals.

Lesson 7 · Observability

Logs, metrics, traces, SLOs, on-call

You cannot fix what you cannot see

LLM backends fail in partial ways: provider 429s, slow embeddings, silent empty outputs, queue lag, wrong tenant metrics. Logs, metrics, and traces are how you prove what happened. Interviewers ask how you would know production is broken before Twitter does — answer with SLIs/SLOs and alerts, not “we check logs sometimes.”
Three pillars: metrics for cheap aggregation and alerts; logs for event detail; traces for request path across services. Correlate with request_id / trace id on every line. Without correlation, on-call is archaeology.

RED and USE methods

For request services: RED — Rate, Errors, Duration. For resources (CPU, queue, pool): USE — Utilization, Saturation, Errors. Build dashboards that answer: is it broken, how bad, since when, which dependency?
text
1SLI examples (LLM API product)2  availability:      proportion of non-5xx on /v1/runs3  latency:          p99 start-to-202 or time-to-first-token4  freshness/lag:    queue oldest-message age < 60s5  quality proxy:    empty-output rate, tool-error rate6  cost:             tokens per successful run (guardrail)78SLO example9  99.9% monthly availability on control plane APIs10  p99 < 300ms for GET /runs/{id}11  async run start lag p95 < 30s1213Error budget: 1 - SLO. Spend budget on shipping; freeze when empty.

Structured logging

JSON logs with stable fields: timestamp, level, service, request_id, tenant_id, run_id, error_code. Do not log secrets, raw PII prompts, or API keys. For LLM apps, log model, token counts, latency, finish_reason — not necessarily full prompt text in default prod.

Metrics cardinality

Labels must be low cardinality: route template not raw URL, status class, model name — not user_id or full prompt hash on every metric. High cardinality metrics will melt Prometheus-style systems and your bill.

Distributed tracing

Propagate trace context across API → queue → worker → provider HTTP. Spans: auth, db, redis, provider call. When p99 spikes, traces show which span moved. Instrument queue publish/consume carefully so async work still links to the originating request when possible.
json
1# Minimal correlation fields on every log line2{3	"ts": "2026-07-15T12:00:00Z",4	"level": "info",5	"service": "run-worker",6	"trace_id": "a1b2…",7	"request_id": "req_…",8	"tenant_id": "ten_…",9	"run_id": "run_…",10	"event": "provider_call_done",11	"model": "…",12	"latency_ms": 842,13	"input_tokens": 1200,14	"output_tokens": 400,15	"http_status": 20016}

Alerting that humans can survive

Alert on symptoms (SLO burn, lag, error rate) more than causes (CPU). Page severity must match user pain. Include runbook links. Suppress flapping with windows. If every alert is ignored, you have no alerts — only noise.

On-call and incidents

Incident loop: detect → mitigate (rollback, flag, shed load) → communicate → root cause → prevent. For LLM spend incidents: kill switches on providers, per-tenant budgets, anomaly alerts on tokens/hour. Write the postmortem with system fixes, not hero worship.
Severity rubric: SEV1 user-facing outage or data leak; SEV2 major degradation/partial outage; SEV3 limited impact. Match response (all-hands vs next business day) to severity. LLM cost blowouts can be SEV1 financially even if “the API returns 200” — define that in advance.

Dashboards that answer questions

Top of dashboard: is user traffic healthy? Then dependency health (DB, Redis, queue, provider). Then tenancy breakdown for noisy neighbors. Avoid 40-panel vanity boards. Every panel should map to a decision: scale, rollback, page provider, or kill a tenant runaway.

Sampling and continuous profiling

Trace sampling (e.g. 1–10% success, 100% errors) keeps cost sane. Tail-based sampling keeps slow traces. Profiling (py-spy, pprof) finds CPU hotspots metrics only hint at. Interviewers like hearing you balance observability cost against debug power — unbounded retention of full traces is a bill.

Synthetic checks vs real traffic

Synthetics hit /health and a cheap authenticated probe continuously. They catch total outages and cert expiry. They miss rare tenant-specific authz bugs — pair with RUM/server SLIs on real traffic. Health endpoints must not be “select 1” only if the product also needs Redis and the queue to work; define deep vs shallow health carefully for k8s probes.
Shallow health: process up. Deep health: can read DB, ping Redis, publish/consume test queue (carefully). Kubernetes liveness should stay shallow so a dependency blip does not kill all pods in a restart loop; readiness can be deeper so traffic shifts away.

Cost of observability itself

Log volume, metric series, and trace retention are budget lines. Set retention tiers (7d hot traces, 30d metrics, longer for audit logs). Sample success paths; keep errors. A blank check for “store everything forever” becomes the next finance incident after the LLM bill.
text
1ON-CALL STARTER PACK (LLM backend)2Dashboard: API RED | queue lag | provider errors | spend/hour3Pages:     SLO burn multi-window | lag > threshold | DLQ depth4          | spend anomaly (z-score or budget)5Runbooks:  provider outage, poison message, key leak, cost spike6Kill:      feature flag stop_new_runs; disable model; revoke key7Comms:     status page template with customer-safe wording

Interview answers — observability

  1. 01Q: Three pillars? Metrics for aggregates/alerts, logs for detail, traces for cross-service latency — correlated by ids.
  2. 02Q: SLI vs SLO vs SLA? SLI measured; SLO target we commit internally; SLA contractual with consequences.
  3. 03Q: What do you alert on? SLO burn / user symptoms first; avoid raw CPU pages without user impact.
  4. 04Q: Cardinality? No user_id labels on metrics; use bounded enums; exemplars/traces for drill-down.
  5. 05Q: Log PII? Default no; redact; separate secure stores for needed audit with access control.
  6. 06Q: Async traces? Propagate trace/run ids in message attributes; worker continues the trace.
  7. 07Q: Golden signals? Latency, traffic, errors, saturation (Google SRE) — map to your RED/USE dashboards.
  8. 08Q: LLM-specific? Token rates, provider errors, empty outputs, tool failures, cost per tenant, queue lag.
  9. 09Q: Error budget? Allowed unreliability; spend to ship features; freeze risky deploys when exhausted.
  10. 10Q: First five minutes of an incident? Confirm impact, mitigate, stop the bleeding, then dig for cause with traces/logs.
docsGoogle SRE Book — Monitoring Distributed SystemsGoogle SREdocsGoogle SRE Workbook — Alerting on SLOsGoogle SREdocsOpenTelemetry DocumentationOpenTelemetrydocsPrometheus — Metric and label guidancePrometheus

Checkpoint

Which is the best primary paging signal for an API that creates LLM runs?

ACPU utilization on a single worker host exceeding 70%.BSLO burn on create-run success rate / latency and async start lag exceeding budget policies.CNumber of lines written to debug logs per minute.
Sign up free to answer and see why

Checkpoint

Why is labeling a Prometheus counter with raw user_id dangerous?

APrometheus forbids string labels entirely by specification.BUnbounded cardinality explodes time series, memory, and bill — use bounded labels and traces for per-user debug.CUser ids are always integers so labels cannot store them.
Sign up free to answer and see why

Checkpoint

A request fails in the worker after the API returned 202. How do you debug end-to-end?

AIgnore the API side; workers are a separate product with no shared identifiers.BPropagate run_id and trace context from API enqueue into the message; worker logs/spans continue the chain.COnly use screenshots of the UI error toast as the observability strategy.
Sign up free to answer and see why

Checkpoint

What belongs in default production logs for an LLM provider call?

AFull prompt text, API keys, and stack traces on success paths.BModel, latency, token counts, status/finish_reason, run_id/tenant_id, error code — not secrets or default full prompts.CNothing — logging is illegal in SaaS backends.
Sign up free to answer and see why

Checkpoint

Error budget is exhausted mid-month. What is the reliability-minded product response?

AIgnore it — SLOs are documentation, not decision tools.BSlow risky feature launches; focus on reliability work until budget recovers per policy.CDelete all dashboards so the budget cannot be measured as empty.
Sign up free to answer and see why

Can you define SLIs/SLOs, a dashboard, and a page-worthy alert for an LLM run pipeline?

New to itGetting thereConfident

Takeaways

  • Metrics, logs, traces — correlated by request/run/trace ids.
  • Alert on SLO symptoms; keep metric cardinality bounded.
  • Structure logs for ops; redact secrets and default PII.
  • Error budgets turn reliability into shipping decisions.

Next: capstone — design a reliable LLM feature backend end-to-end.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.