Lesson 3 of 6 · 46 min

Observability & feedback loops

The discipline that lets you tell a customer in a Friday-night incident review what actually happened — the LLM span schema, why time-to-first-token is the new p95 (and how Intercom cut 2s with it), feedback loops that close the system, and continuous-eval drift detection.

What lets you answer “what actually happened?”

Observability is the thing that lets you tell a customer, in a Friday-evening incident review, exactly what the system did. Honeycomb defines LLM observability as “the discipline of monitoring, tracing, and analyzing every stage of using an LLM in production” — granular insight into how the model behaved so you can troubleshoot faster and improve quality in real time. For a solutions engineer this is both the production-AI-health round in an interview and the artifact that turns “the bot is dumb” into a precise, span-localized diagnosis in front of the customer.
The anti-pattern observability kills: aggregates hide the problem. An average token count or a single p95 latency number tells you nothing about which prompt variant, which retrieval slice, or which tenant drove a cost or quality leak. Honeycomb’s production guidance is explicit that high-granularity events — prompt variant, retrieval context, token counts, specific user IDs, RAG chunk injection — are what let you slice telemetry by model configuration, user cohort, or prompt version to find the leak. You instrument the span, not the dashboard average.
The minimum LLM span schema is well understood, and reciting it is high-signal: model identifier, prompt version, prompt-template hash, dataset/version fingerprint, retrieval context (chunk IDs and scores), tool calls (name, args, latency, success), token counts (input / output / cached / total), and latency cohorts (time-to-first-token vs total). With those dimensions on every span you can answer “which prompt version regressed for which tenant on which retrieval slice” without redeploying instrumentation.
code
1MINIMUM LLM SPAN SCHEMA (emit on every model call)23  model.id              gpt-4o-2024-XX        which model + version4  prompt.version        v37                   bisect a regression to a change5  prompt.template_hash  ab12cd                detect silent prompt edits6  retrieval.chunk_ids   [c19, c8, c204]       was the gold chunk even fetched7  retrieval.scores      [0.81, 0.79, 0.62]    rank quality of the context8  tool.calls            [{name,args,ms,ok}]   which tool, how long, success9  tokens                {in, out, cached, total}  cost + the cache hit-rate10  latency               {ttft_ms, total_ms}   the TWO latencies, separately11  tenant / user / cohort                       slice by who hit it1213  Aggregates hide leaks. Slice by prompt.version x tenant x retrieval slice.
LLM Observability, Evaluation, Experimentation PlatformAI Engineer (Dat Ngo, Arize)

Trace the agent, not just the call

For a single LLM call, a span is enough; for an agent that plans, retrieves, calls three tools, and reflects, you need a distributed trace — a tree of spans with one trace id, so you can see the whole trajectory and find which step failed, not just that the final answer was wrong. This is the difference between outcome analytics (the answer was bad) and trajectory analytics (the retrieval step returned the wrong chunk, so the second tool call was made on bad input). The agents-towards-production playbook treats monitoring as a non-skippable phase precisely because a multi-step agent has many more places to fail silently than a single call.
The mechanism is OpenTelemetry-style context propagation: the orchestrator opens a root span, every retrieval/model/tool call is a child span carrying the same trace id plus its LLM attributes, and the tree reconstructs the agent’s decision path. Honeycomb ships an Agent Timeline view and a Honeycomb MCP server so an agent can even query its own traces — but the portable lesson is vendor-agnostic: emit OTel spans with LLM semantic attributes and you can slice the trajectory in any backend (Honeycomb, Arize Phoenix, Langfuse). Interview angle. “Your agent gives a wrong final answer — how do you find where it went wrong?” → pull the distributed trace, walk the span tree, and localize: was it a retrieval miss (wrong chunk ids), a bad tool result (tool span shows the error), or a reasoning slip (the right inputs were present)? Naming the trace tree, not “check the logs,” is the senior signal.
code
1OUTCOME vs TRAJECTORY (one trace id, a tree of spans)23  trace: reroute-9f24   |- orchestrator.plan            120 ms5   |- retrieval.search             40 ms   chunk_ids=[c8,c19]  scores=[.81,.62]6   |- tool.tms_lookup              900 ms  ok=true             <- the tail!7   |- model.decide                 600 ms  tokens={in:1.2k,out:90}8   |- tool.carrier_api.reroute     err: 429 rate-limited       <- the FAILURE9   `- model.summarize             300 ms1011  Outcome view: "the reroute failed." Trajectory view: "carrier_api 429'd."12  You can only fix what the SPAN TREE localizes -- not what the aggregate hides.

Time-to-first-token is the new p95

The most important latency insight for a customer-facing LLM: model a response as three sequential spans — time-to-first-token (TTFT), generation, and post-processing — so you can attribute user-perceived latency to its dominant cause. Honeycomb’s Intercom case study is the proof: Intercom wrapped TTFT as a discrete span around the LLM call for its Fin.ai agent and reduced median response times by two seconds while gaining per-token cost visibility. Without splitting TTFT from total, you cannot tell whether the user is waiting on prompt prefill or on a long generation — and they have different fixes (shrink/cache the prompt vs shrink/stream the output).
Interview angle. “Users say the assistant feels slow — how do you find out why?” The senior answer is span-level: instrument TTFT separately from total latency, slice by prompt version and retrieval slice, and check whether the tail is prefill (big prompt → cache it), generation (long output → cap/stream it), or a slow tool call (the span shows it). Quoting Intercom’s two-second win signals you optimize the tail at the span level, not the aggregate. The same telemetry is also how you detect context truncation (a span shows the prompt hit the window cap) and attack traffic (prompt-injection attempts show up as anomalous tool calls) — both are observability problems in disguise, per Honeycomb’s “hard stuff nobody talks about” lessons.
Latency is perceived as quality. The team that instruments the three phases separately — and shrinks whichever one is the tail — wins on speed and cost at once; the team that watches a single p95 ships a slow assistant and never knows which knob to turn.

Feedback loops close the system

Telemetry tells you what happened; feedback loops turn that into improvement. The pattern, drawn from production feedback-loop guides: capture explicit and implicit signals and attach them to the trace (thumbs, copy-to-clipboard, “regenerate,” dwell time, downstream task success), annotate production traces with that signal, classify the annotated cohort with an LLM to bucket feedback into failure modes, prioritize fixes, then push the fixes back into the eval harness as new golden-set cases. Microsoft reinforces the key constraint: thumbs alone are insufficient; pair them with structured, failure-mode-specific capture.
code
1THE FEEDBACK LOOP (telemetry -> evals, closed)23  1. CAPTURE    attach signals to the trace4                explicit: thumbs, regenerate, edit, copy5                implicit: dwell time, downstream task success/failure6  2. ANNOTATE   tag production traces with the signal7  3. CLASSIFY   LLM buckets the cohort into FAILURE MODES8                (wrong citation / refused tool / stale data / hallucination)9  4. PRIORITIZE rank by frequency x severity x effort10  5. FEED BACK  failing cases become NEW golden-set eval cases (Lesson 2)1112  A thumbs button is one bit. Instrument the specific failure modes.
Two practical details that separate a real loop from a backlog of thumbs-downs. First, implicit signals often beat explicit ones: only a few percent of users ever click a thumb, but nearly all of them reveal satisfaction through behavior — did they copy the answer, immediately rephrase and retry, abandon the session, or complete the downstream task? Attaching downstream task success to the trace is the highest-signal feedback you can collect, and it requires no UI button. Second, the classify step is itself an LLM job: run a model over the annotated cohort to bucket free-text complaints and implicit signals into the same named failure modes your evals test (wrong citation, refused tool, stale data, hallucination), then rank by frequency × severity. That ranked list is your eval backlog and your roadmap in one artifact.
For a solutions engineer this loop is also a customer-facing asset: the trend line of a falling wrong-citation rate, or a rising task-success rate, is exactly the evidence a champion needs to defend the renewal in front of the steering committee. The observability you build to debug the system doubles as the proof-of-value you bring to the QBR — which is why “how will we see it improving over time?” is a discovery question worth pre-answering with this loop, tying Lesson 3 straight back to Lesson 5’s value story.

Drift detection is the eval harness’s older sibling

Model and retrieval performance drift whenever the data, the prompt, the upstream API, or the user population changes — and none of those throw an error. The fix is continuous eval: a scheduled job that re-runs a representative slice of the golden dataset against production traffic and alerts on regression in quality, latency, and cost. Anthropic’s framework treats drift-catching regression suites as a permanent fixture, not a one-time gate. This is where Lessons 2 and 3 fuse: the eval suite you built offline becomes a monitor when you schedule it against live traffic and wire its regressions into the same alerting as your latency SLOs.
Two named production lessons sharpen the point. Honeycomb’s “All the Hard Stuff Nobody Talks About” concludes the LLM should be viewed as an “engine for features,” not a standalone product — which means you instrument the feature’s outcome, not just the model’s output. And the silent-regression failure mode from earlier in the track applies here: a provider can update a model under you, or three small changes can interact (a reasoning-effort cut, a caching bug, a prompt tweak) and read as broad quality loss; only pinned versions plus a continuous eval catch it before the customer does. Interview angle. “The customer says it got worse this week but you shipped nothing — what do you check?” → prompt-version and model-version telemetry, a continuous-eval regression alert, and the feedback cohort; a sustained drop with no deploy is the classic provider-side or interacting-change signature.
code
1WHAT DRIFTS, AND HOW THE MONITOR CATCHES IT23  Drift source            Symptom                  Detection signal4  ---------------------   ----------------------   -------------------------5  Provider model update   broad quality drop,      pinned model.id + eval6                          no deploy                regression alert7  Prompt edit             one path regresses       prompt.version diff in spans8  Upstream API change     a tool starts failing    tool-call success rate drop9  User population shift    new query types miss     online sampling finds the10                          the golden set            unseen cohort11  Retrieval staleness     citations go stale       freshness timestamp + recall1213  The monitor is the offline golden set RE-RUN on a schedule against live14  traffic -- quality, latency, AND cost regressions alerted like an SLO.
articleAI in Production Is Growing Faster Than We Can Trust It (Intercom Fin.ai, TTFT)HoneycombarticleAll the Hard Stuff Nobody Talks About when Building LLM ProductsHoneycombdocsWhat Is LLM Observability and Monitoring?HoneycombarticleUser Feedback Loops: Closing the AI Data CycleFuture AGI

Checkpoint

A customer complains the assistant “feels slow.” Your aggregate p95 looks fine. What’s the right diagnostic move?

ATell them p95 is within SLO and the perception is subjectiveBSplit the response into TTFT / generation / post-processing spans and slice by prompt version, tenant, and retrieval slice to locate the tailCUpgrade everyone to a faster model to be safe
Sign up free to answer and see why

Checkpoint

You can ship only ONE additional latency signal on every LLM span. Which buys the most diagnostic power for a customer-facing agent?

AA single end-to-end total-latency numberBCPU utilization of the inference hostCTime-to-first-token, recorded separately from total latency
Sign up free to answer and see why

Checkpoint

Your team keeps re-discovering the same failure (wrong citations) every quarter. What’s missing from the system?

AA closed feedback loop: capture signals on traces → classify into failure modes → fix the top bucket → convert fixes into permanent eval casesBA bigger model with better world knowledgeCMore thumbs-up / thumbs-down buttons in the UI
Sign up free to answer and see why

Checkpoint

The customer says quality dropped this week, but your team shipped nothing. What do you check first?

AAssume random LLM variance and wait to see if it recoversBPrompt-version and model-version telemetry plus a continuous-eval regression alert and the feedback cohort — a no-deploy drop is the provider-side or interacting-change signatureCRoll back the most recent infrastructure change
Sign up free to answer and see why

Checkpoint

You’re instrumenting a RAG assistant and must choose what to attach to each span for the most useful cost analysis. Best choice?

AOnly the total token count per requestBInput / output / cached / total token counts plus prompt version, tenant, and retrieval slice as dimensionsCA daily dollar total emailed to the team
Sign up free to answer and see why

Interview prep

The observability round tests whether you instrument at the span level, separate the two latencies, attach cost and cohort dimensions, close the feedback loop, and run continuous eval for drift. Strong answers name the span schema and the Intercom TTFT win and treat the LLM as an “engine for features” whose outcome you measure; weak answers stop at aggregate dashboards and a thumbs button.
  1. 01“What do you put on an LLM span?” → model id, prompt version + hash, retrieval chunk ids/scores, tool calls, token counts (in/out/cached/total), TTFT + total, tenant/cohort.
  2. 02“Why TTFT separately from total?” → it splits prefill from generation, which have different fixes; Intercom cut median response time 2s by wrapping it.
  3. 03“The assistant feels slow — how do you find why?” → three spans + slice by prompt version/tenant/retrieval slice; check prefill vs generation vs tool tail.
  4. 04“How do you close the loop?” → capture multi-bit signals on traces → classify failure modes with an LLM → fix top bucket → feed fixes back as golden-set cases.
  5. 05“Why aren’t thumbs enough?” → one-bit sentiment, poor correlation with task success; instrument specific failure modes (wrong citation, refused tool, stale data).
  6. 06“How do you detect drift?” → scheduled continuous eval on a golden slice vs live traffic, alerting on quality/latency/cost regressions like an SLO.
  7. 07“It got worse with no deploy — what do you check?” → version telemetry + continuous-eval alert + feedback cohort; classic provider-side/interacting-change signature.
  8. 08“How do you catch a prompt-injection attempt in prod?” → it surfaces as anomalous tool calls / context anomalies in the spans — an observability problem in disguise.
Going deeper, the follow-ups probe how telemetry composes with the rest of the stack: “how do you attribute a cost spike to a root cause?” (slice token counts by prompt version × tenant × retrieval slice — the aggregate can’t); “how long did a hallucinated answer live before you caught it?” (detection latency — only online eval + feedback answers it, and it’s a number the customer will ask for); and “what’s the difference between monitoring the model and monitoring the feature?” (Honeycomb’s “engine for features” point — you instrument the user-facing outcome, with the model span as one contributor). Tie each back to a span dimension.

Could you design the span schema, justify TTFT instrumentation with the Intercom number, and describe the feedback loop that feeds evals?

New to itGetting thereConfident

Takeaways

  • Aggregates hide leaks — instrument the span, slice by prompt version × tenant × retrieval slice.
  • TTFT is the new p95: split it from total latency; Intercom cut 2s of median latency by wrapping it.
  • Emit the full LLM span schema so a regression can be bisected to a prompt/model version.
  • Close the loop: capture multi-bit signals → classify failure modes → feed fixes back into evals.
  • Thumbs are one bit; instrument specific failure modes for real signal.
  • Continuous eval is drift detection — schedule the golden set against live traffic and alert like an SLO.

Next: how the AI reaches the customer’s systems at all — MCP, enterprise connectors, and the N×M governance problem.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.