“Works in staging, fails in prod” is an observability gap. Span-level traces of every action/tool/cost/failure, cost as the earliest regression signal, the online/offline eval loop, and shadow/canary rollout — how teams debug agents at scale, and the interview round that probes it.
Agents that pass every staging test still fail in production — because staging uses developer-curated inputs with predictable tool sequences, while production pulls untrusted content (customer PDFs, web pages, emails) that perturbs trajectories in ways your eval suite never saw. Traditional APM tells you the API responded with low latency and no 500s; for an agent you need to know why it chose tool A over tool B. One widely-cited figure: ~46% of agent POCs fail on exactly this observability gap.
The mental model to internalise: observability for agents is decision-level, not service-level — your unit of debugging is the trajectory, not the HTTP request. Why is the staging/prod gap so much wider for agents than for ordinary services? Because an agent’s behaviour is a function of its inputs at every step, and production inputs are adversarial and long-tailed in ways curated test inputs are not. A customer PDF with an unusual table, a web page with a hidden instruction, an email thread three times longer than anything in your fixtures — each perturbs the trajectory, and the agent’s autonomy means a single perturbed observation can send it down a path your evals never exercised. Interview angle. “Why do agents that pass staging fail in prod?” is testing whether you understand that the input distribution, not the code, is what changed.
Span-level tracing: the unit of agent debugging
Instrument every operation as a span — generation, tool call, retrieval, event — capturing the prompt, the response, token usage, latency, cost, and the exact arguments. Langfuse (acquired by ClickHouse, Jan 2026) and Braintrust are the OpenTelemetry-aligned backbones; their value is letting you drill from a failed task score straight to the specific span that caused it (Braintrust renders the run as an expandable trace tree). Anthropic caught a Claude Code bug — internal NaNs in tool-call arguments stalling CORE-Bench at 42% — only by grading the trajectory spans; fixing it took the score to 95%.
The structure that makes this work is a trace = a tree of spans, where each span has a parent, a start/end time, and structured attributes. A single agent task is one trace; its generations and tool calls are child spans; a sub-agent (if any) is a nested sub-tree. OpenTelemetry’s emerging GenAI semantic conventions standardise the attribute names (model, token counts, tool names) so your traces are portable across backends rather than locked to one vendor’s SDK. The reason to align to OTel rather than hand-roll: you almost certainly already run distributed tracing for your services, and an agent trace that lives in the same system lets you correlate “the agent looped” with “the downstream API was slow” in one view.
python
1with tracer.span("agent.run", input=goal) as root:2 for step in range(max_steps):3 with tracer.span("generation", parent=root) as g:4 resp = model(messages, tools=TOOLS)5 g.set(tokens=resp.usage, cost=resp.cost, latency_ms=resp.ms)6 if resp.tool_call is None:7 break8 with tracer.span("tool", parent=root) as t: # one span per tool call9 t.set(name=resp.tool_call.name, args=resp.tool_call.args)10 obs = call(resp.tool_call)11 t.set(ok=obs["ok"], error=obs.get("error"))12# now a failed task score links to the exact tool span that broke it
What to put on a span (and what not to)
A span is only as useful as its attributes. The load-bearing fields: inputs/outputs (the prompt and the model’s message — redacted), token usage and cost per call, latency split into TTFT and total, the tool name and arguments, the structured error if any (the TRANSIENT/PERMANENT/REQUIRES_HUMAN tag from L1), and a session/user id so you can reconstruct a multi-turn conversation. The scale gotchas that bite later: traces of agent runs are large (full prompts re-sent every step), so you sample (keep 100% of failures and a fraction of successes) and set retention; and prompts contain PII, so you redact at the SDK before it leaves your process, or your observability store becomes a compliance liability. A trace you can’t store cheaply or can’t legally keep is not observability.
Online vs offline evals: the production feedback loop
Observability and evals are one loop, not two disciplines. Offline evals run your regression suite against fixed datasets in CI (lesson 4). Online evals run graders — cheap deterministic checks, and a sampled LLM judge — against live production traces, continuously. The point of online evals is twofold: they score behaviour on the real, drifting input distribution that offline sets miss, and the failures they surface are exactly the traces you promote into the offline regression set. That promotion is the flywheel: production reveals a new failure mode → you capture the trace → it becomes a permanent test → the next change that reintroduces it fails CI. Interview angle. “How do offline and online evals fit together?” → offline gates changes pre-ship; online watches the live distribution and feeds new failures back into the offline set — a closed loop, not a one-time test.
Cost is your earliest regression signal
The most reliable early warning of a prompt-template regression isn’t a quality drop — it’s a cost jump. A “small” tweak that nudges the agent to make one extra tool call shows up in tokens-per-task before it shows up in answer quality. So emit decision-aware events (“agent selected tool X with args Y”) to replay the A/B decision diff, and alert on cost-per-resolved-task, not cost-per-call.
The mechanism is the super-linear cost growth from L1: an extra tool call doesn’t add one call’s worth of tokens, it adds that observation plus the re-sent transcript on every subsequent step. So a one-step regression can move cost-per-task by far more than 1/N, which is precisely why cost is a sensitive leading indicator — it amplifies the change. The metric discipline matters: cost-per-call can look flat while cost-per-resolved-task climbs (the agent is making more calls but resolving the same number of tasks), and cost-per-resolved-task is the one tied to business value. Pair it with steps-per-task and tool-call distribution dashboards; a shift in which tools get called, or how many, is the fingerprint of a prompt or model change before any quality metric moves.
code
1WHAT TO ALARM ON (leading -> lagging)23 signal why it moves first alert on4 ------------------------ ----------------------------- --------------------5 cost-per-resolved-task extra step re-sends transcript step-change vs 7d median6 steps-per-task (p50/p95) regression adds tool calls p95 creep7 tool-call distribution prompt/model shifts tool choice KL-divergence from base8 cost-per-CALL (alone) MISLEADING -- can stay flat do NOT rely on this9 final-answer quality LAGGING -- users already hit it still track, but late10 HTTP 500 / latency (APM) infra only -- agent stays up necessary, not sufficient
Shadow mode and canary: catch drift that isn’t in your git log
Before full rollout, run the agent in shadow mode in parallel with humans (a ~30-day window is common), building a golden set from the real failures you observe, then canary a small slice. The drift that bites most isn’t a change you made — a provider silently updates a model checkpoint and quality sags weeks later (Anthropic’s own postmortem traced two months of complaints to interacting changes). Per-span traces measured on the production distribution are what surface it.
The rollout ladder is worth naming in order, because each rung answers a question the previous one couldn’t. Shadow mode: the agent runs on real inputs but its outputs are logged, not served — you measure behaviour on the production distribution with zero user risk, and you mine the disagreements with the human/incumbent for your golden set. Canary: serve a small slice (1-5%) and compare cost/quality/steps against the control on the same traffic. A/B: a controlled split to measure the business metric. Full rollout with the online-eval alarms above standing watch. The Anthropic postmortem is the cautionary tale that makes this concrete: three small interacting changes (a reasoning-effort cut, a caching bug dropping older thinking, a verbosity tweak) read as months of broad “it got dumber,” and only production-distribution traces could have isolated them — no single git commit was the culprit.
Interview prep
Observability interviews test whether you can debug an agent you can’t see. The signature question is “works in staging, fails in prod — what now?”, and interviewers want decision-level traces, the input-distribution insight, the online/offline eval loop, and a staged rollout — not “add retries” or “upgrade the model.” The other favourite is the leading-indicator question: name cost-per-resolved-task before quality, and explain why (super-linear cost amplifies a one-step regression).
01“Why do agents that pass staging fail in prod?” → The input distribution changed: untrusted, long-tailed production content perturbs trajectories your curated tests never exercised. Fix with real-trace visibility + staged rollout.
02“What’s the unit of agent debugging?” → A span: instrument every generation, tool call, and retrieval with prompt, tokens, cost, latency, args, and structured error — a trace is a tree of spans, OTel-aligned.
03“Earliest signal a prompt change regressed the agent?” → A jump in cost-per-resolved-task / steps-per-task — an extra tool call re-sends the transcript and shows up in cost before quality.
04“Why cost-per-resolved-task, not cost-per-call?” → Cost-per-call can stay flat while the agent makes more calls per task; only cost-per-resolved-task tracks business value and reveals the regression.
05“How do online and offline evals fit together?” → Offline gates changes in CI; online grades live traces on the real distribution; online failures get promoted into the offline regression set — a closed loop.
06“How do you roll out a new agent version safely?” → Shadow mode (log, don’t serve) → canary (1-5%) → A/B → full rollout, with online-eval alarms watching cost, steps, and tool distribution.
07“What do you NOT trust from APM for an agent?” → A green APM dashboard (uptime, latency, 500s) is consistent with a badly regressed agent; APM is service-level, agents need decision-level.
08“What are the scale gotchas of agent tracing?” → Traces are large (sample: keep all failures, a fraction of successes), and prompts carry PII (redact at the SDK) — or observability becomes a cost and compliance liability.
Going deeper. Reach for the Anthropic postmortem as a case study of why git-log debugging fails: multiple small interacting changes produced a large, diffuse quality complaint that only production-distribution traces could isolate. Mention OTel GenAI semantic conventions (portable traces, correlate with existing service tracing), tool-call-distribution drift measured as KL-divergence from a baseline (a fingerprint of a model/prompt change), and the NaN-in-tool-args Claude Code bug (CORE-Bench stuck at 42%, fixed to 95%) as proof that trajectory spans — not output scores — are what catch a whole class of mechanical bugs.
Your agent passes every staging test but misbehaves in production. Most likely cause and fix?
AThe model is too small — upgrade itBProduction feeds untrusted content that perturbs trajectories; add span-level traces and a shadow-mode rollout on real trafficCAdd more retries
Cost-per-call on your agent has been flat for weeks, yet the monthly bill is climbing and tasks resolve at the same rate. What’s happening?
ACost-per-call is flat, so nothing changed — the finance team miscountedBThe agent is making more tool calls per resolved task (cost-per-resolved-task and steps-per-task are up); a change nudged extra calls and the re-sent transcript amplified the costCOutput tokens got more expensive
You’re about to ship a new agent version and want to measure its real-world behaviour with zero user risk before serving any of its output. Which rollout stage?
ACanary at 5% of live trafficBA full A/B testCShadow mode: run the agent on real inputs but log its outputs instead of serving them, then mine disagreements with the incumbent for your golden set
Your team wants to keep full prompt/response traces of every production agent run forever for debugging. What’s the senior concern?
ATraces are large (full transcripts re-sent every step) and contain PII — sample (keep all failures, a fraction of successes), set retention, and redact at the SDK before storageBNo concern — storage is cheap, keep everythingCOnly keep the final answers to save space
Could you instrument span-level traces, run the online/offline eval loop, alert on cost-per-task, run a shadow/canary rollout, and defend it in an interview?
Not yetMostlyConfident
Takeaways
“Works in staging, fails in prod” is a decision-level observability gap — untrusted, long-tailed prod inputs perturb trajectories your tests never saw.
Trace every generation/tool/retrieval as a span (a trace is a tree of spans, OTel-aligned) with cost, latency, args, and structured error.
Cost-per-resolved-task and steps-per-task are the earliest regression alarms (super-linear cost amplifies an extra step); cost-per-call alone misleads.
Run the online/offline eval loop: online grades live traces and feeds new failures into the offline regression set.
Ship behind shadow mode → canary → A/B; the worst drift is a silent provider model update (Anthropic’s postmortem).
Sample large traces and redact PII at the SDK, or observability becomes a cost and compliance liability.
Next: securing tool-using agents — prompt injection, the lethal trifecta, and gating writes.