Lesson 6 of 6 · 52 min

Capstone: a summarize-and-extract microservice

Assemble the whole track — prompt contract + structured output + reliable client + caching + tests + observability — into a service that turns a document into a structured summary you can trust in production, and that you can defend end-to-end in a system-design interview.

Capstone

A summarize-and-extract microservice

Where the five lessons become one service

Each lesson so far solved one problem in isolation. The capstone is the integration test: a real service that takes a document and returns a structured summary plus extracted fields (entities, action items, owner), behind a reliable client, with caching, optional streaming, an observability span, and tests that work on probabilistic output. This is the canonical “ship your first LLM feature” build — and it’s also the exact thing a system-design interviewer asks you to whiteboard, so we’ll build it and rehearse defending every decision.

The architecture

Request → assemble the prompt (a stable, cached system contract + just-in-time context, L2) → call through the reliable client (timeout / classified retry+jitter / fallback / breaker, L4) → strict structured output into a Pydantic schema (L3) → validate (and repair on business-rule failures only) → cache (provider prompt cache always; semantic cache if traffic repeats, L5) → return, logging a token/cost/latency span for every request (L5). Each arrow is a lesson; the service is the composition.
code
1REQUEST PATH (each stage = one lesson)23  doc in4    -> assemble prompt        L2  stable cached system contract + JIT context5    -> semantic cache lookup  L5  hit? return cached structured result, skip the model6    -> reliable client call   L4  timeout -> retry(jitter) -> fallback -> breaker7    -> strict structured out  L3  constrained decoding into Pydantic (no parse step)8    -> business-rule validate L3  validators; targeted re-ask on semantic failure only9    -> write caches + span    L5  provider prompt cache always; log tokens/cost/latency10  structured result out1112  Cross-cutting: pinned model+prompt VERSION (L1/L2) + golden-set eval gate in CI
Interview angle. When you whiteboard this, narrate the data flow as a sequence of guarantees, not boxes: “the prompt is a cached, versioned contract; the client makes the flaky upstream look reliable; the schema makes the output a contract code can consume; the cache and span make it cheap and observable; the eval gate makes change safe.” Naming the guarantee each layer provides — and which failure it removes — is what distinguishes a senior answer from a box-and-arrow diagram.
python
1from pydantic import BaseModel2import instructor, time3from openai import OpenAI45class DocSummary(BaseModel):6    summary: str7    action_items: list[str]8    owner: str | None910client = instructor.from_openai(OpenAI(timeout=45.0, max_retries=3))   # L4 reliability1112SYSTEM = "Summarize the document and extract action items + owner. Use only the document."  # L2 stable prefix1314def summarize(doc: str) -> tuple[DocSummary, dict]:15    t0 = time.time()16    out = client.chat.completions.create(17        model="gpt-4o-mini",                 # L5: cheap tier; escalate only the hard tail18        response_model=DocSummary,           # L3: strict structured output19        max_retries=2,                       # business-rule retries only20        messages=[{"role": "system", "content": SYSTEM},21                  {"role": "user", "content": doc}],22    )23    span = {"latency_ms": int((time.time()-t0)*1000)}   # L5: log cost/latency per request24    return out, span

Handling a document bigger than the window

The first scale question for this service is the L1 one: what if the document exceeds the context window? Don’t reach for a bigger-context model by reflex — walk the menu and pick by the task. For a summary (a holistic question), map-reduce: summarise each chunk in parallel, then fold the partial summaries until they fit, then do a final pass. For targeted extraction (“who is the owner of the renewal clause”), chunk + retrieve the few relevant passages and run extraction only on those — cheaper and more accurate than stuffing. Note the failure mode you’re avoiding: map-reduce can lose cross-chunk context (a fact split across two chunks), so for extraction that spans the whole doc, prefer retrieval with overlap.
python
1def summarize_large(doc: str, chunk) -> DocSummary:2    parts = [summarize(c)[0].summary for c in chunk(doc)]    # MAP: per-chunk (parallel)3    while n_tokens("\n".join(parts)) > WINDOW_BUDGET:4        parts = [summarize("\n".join(g))[0].summary           # REDUCE: fold partials5                 for g in groups_of(parts, 5)]6    return summarize("\n".join(parts))[0]                     # final pass over the fold

Testing a probabilistic service: three concentric rings

You can’t assert byte-equality, so test in rings, cheapest first. Ring 1 — deterministic: schema validation, JSON validity, exact-match on closed-form fields, regex on identifiers (free, run on every request). Ring 2 — semantic: embedding-similarity floors on the summary, refusal-string detection on safety cases. Ring 3 — judge: a stronger model scoring against a rubric, calibrated against a small human-labelled golden set. Ship ring 1 first; add the judge last, and run it on a sampled subset to control cost.
code
1THREE TEST RINGS (cheapest + most reliable first)23  Ring   What it checks                         Cost   Run on4  ----   ------------------------------------   ----   ---------------------5  1 det  schema valid, JSON valid, regex ids,   free   every request + CI6         exact-match closed fields                     (assert these always)7  2 sem  embedding-similarity floor on summary, cheap  CI + sampled prod8         refusal-string on safety prompts9  3 jdg  stronger model scores vs a rubric,     $$     CI gate + small sample10         calibrated to a HUMAN golden set              (sampled to bound cost)1112  Pitfall: an UNcalibrated LLM judge drifts and is biased (length, position).13  Anchor it to ~50 human labels and re-check agreement, or it lies to you.
The eval that actually gates the deploy is a golden set: ~50 production-like documents with known-good expected fields, frozen and version-controlled. CI runs the service against it on every prompt or model change and blocks the merge if any ring regresses past threshold. This is the same discipline as L2’s versioned prompts and L1’s silent-regression defence, made concrete: pin model and prompt_version, score every change, and you convert “users tell you it got worse” into “CI blocks the PR.”

Observability: the span that explains every request

The last 20% that separates a demo from a service is the per-request span: log prompt_version, model, input/output token counts, computed cost, latency (TTFT + total), cache-hit flag, retry/fallback count, and whether validation repaired. With that, every production question becomes answerable: a cost cliff is a trend line on the cost field (L5), a quality dip bisects to a prompt_version (L2), a latency regression splits into prefill vs decode (L1), and an outage shows up as a fallback-rate spike (L4). Without it, you’re debugging a probabilistic black box from user complaints. This is why observability is a capstone requirement, not a nice-to-have.

Case study: the same shape, at scale

This summarize-and-extract shape is exactly what production teams run at volume — support-ticket triage, contract extraction, meeting-notes processing, document classification. The lessons compound the same way they do here, just with bigger numbers: Shopify’s Sidekick (L2) is this service with JIT instructions and an LLM-judge gate; Notion (L1/L5) is this service with a small fine-tuned model on the hot path for latency; Intercom Fin (L4) is this service with SLOs on TTFT and wasted tokens; and Anthropic’s own postmortem (L1) is what happens to this service when you don’t pin versions and eval every change — three small interacting changes read as two months of “it got dumber.” The capstone isn’t a toy; it’s the minimal honest version of what those teams ship.
A production LLM feature is not “a model that answers.” It’s a contract (schema), a shield (reliable client), a budget (cost/latency levers + span), and a gate (golden-set eval on pinned versions). Build those four and the model is just the easy part in the middle.

Interview prep

The capstone maps to the “design an LLM-powered feature/service” system-design round — the most common senior AI-engineering interview. Interviewers want to see you compose the whole stack, justify each layer by the failure it removes, and reach for the right tool under follow-up pressure. Drive the conversation as the request path above, naming the guarantee at each hop.
  1. 01“Design a document summarize-and-extract service.” → cached prompt contract → reliable client → strict structured output → validate/repair → cache + per-request span → golden-set eval gate on pinned versions.
  2. 02“How do you guarantee parseable output?” → strict structured outputs (constrained decoding into a Pydantic schema), not prompt-and-pray + try/except — removes the parse-failure class.
  3. 03“How do you test something nondeterministic?” → three rings: deterministic (schema/exact/regex), semantic (embedding floor), LLM-judge calibrated to ~50 human labels; ship ring 1 first, judge last and sampled.
  4. 04“What gates a deploy?” → a frozen golden set in CI scoring every prompt/model change against pinned model+prompt_version; block the merge on regression.
  5. 05“The doc is bigger than the window — what do you do?” → map-reduce for holistic summaries, chunk+retrieve for targeted extraction; not “bigger-context model” by reflex.
  6. 06“What do you log per request?” → prompt_version, model, in/out tokens, cost, TTFT+total latency, cache-hit, retry/fallback count — so every prod question is answerable.
  7. 07“Users say it got worse but you shipped nothing — what protected you?” → pinned versions + golden-set eval (Anthropic’s postmortem) turn a silent provider regression into a detectable, bisectable event.
  8. 08“How do you make it cheap?” → cache the stable prefix, route the easy majority to a cheap model, keep output short, batch any offline backfill — watch the per-request cost span.
  9. 09“Where does reliability live?” → the L4 stack (timeout → retry+jitter → fallback → breaker) plus idempotency for write-side calls; ideally behind a shared gateway.
  10. 10“What’s your rollout plan for a prompt change?” → canary on a traffic slice, compare on the golden set + live metrics, promote on no regression, rollback is a config flip.
Push it deeper. Expect integration follow-ups that cross lessons: “The output is valid JSON but the owner field is wrong half the time — where in the stack?” (an accuracy/grounding problem, not structure: cite the source span, add a validator, reorder reasoning-before-field, eval it — L3). “Cost doubled at 10× traffic — first place you look?” (the per-request cost span: output-length creep, an un-cached prefix, or a router escalation drift — L5). “Your provider has a partial outage — what does the user see?” (the breaker fails over to a fallback model whose answers you’ve eval’d, and the span’s fallback-rate spikes — L4). “How do you know the LLM judge itself is trustworthy?” (calibrate it against the human golden set and monitor agreement; an uncalibrated judge has length/position bias and drifts).
repoopenai-cookbook — structured outputs, retries, evals, end-to-end examplesopenai/openai-cookbookarticleA practical guide to LLM regression testing (golden datasets)Evidently AIarticleYour AI product needs evals (how to build the golden-set eval loop)Hamel Husain

Checkpoint

You’re writing the first tests for the service’s probabilistic output. Which ring do you ship first?

AAn LLM-as-judge rubric scoring overall qualityBDeterministic checks: schema validation, JSON validity, exact-match on closed fieldsCByte-for-byte snapshot comparison against a saved output
Sign up free to answer and see why

Checkpoint

Two weeks after launch, users report the service “got worse,” though you changed nothing. What would have protected you — and now tells you what happened?

AHigher temperature for more diverse outputsBA pinned model/prompt version plus a CI eval gate on a golden set that runs on every changeCA bigger context window
Sign up free to answer and see why

Checkpoint

A user uploads a 400-page contract — well over the context window — and asks for a one-paragraph summary. Best approach for this service?

ASwitch to the largest-context model and stuff the whole document inBMap-reduce: summarize chunks in parallel, fold the partial summaries until they fit, then a final passCTruncate the contract to the first N pages that fit
Sign up free to answer and see why

Checkpoint

Strict structured output is on, so JSON is always valid — yet the extracted owner field is wrong on ~30% of documents. Where is the problem and the fix?

AIt’s an accuracy/grounding problem, not a structure one — ground the field in a cited source span, add a validator that the span exists, and order a reasoning field before ownerBConstrained decoding is failing — disable strict modeCRaise max_tokens so the model has room to get it right
Sign up free to answer and see why

Checkpoint

You add an LLM-as-judge as ring 3 and trust its scores to gate deploys. A month later it’s passing changes that users dislike. Most likely cause and fix?

AThe judge model is too small — always use the biggest model as judgeBThe judge is uncalibrated and has drifted (length/position bias) — anchor it to a ~50-example human-labelled golden set and monitor judge-vs-human agreementCSwitch the gate to byte-for-byte snapshot matching instead
Sign up free to answer and see why

Could you build, test, observe, and deploy-gate this microservice end-to-end, handle the bigger-than-window and accuracy cases, and defend the whole design in a system-design round?

Not yetMostlyConfident

You can now

  • Assemble prompt contract + structured output + reliable client + caching + observability span into one service.
  • Scale past the context window with map-reduce (summaries) or chunk+retrieve (targeted extraction).
  • Test probabilistic output in three rings and gate deploys on a golden set with a calibrated judge.
  • Pin versions and eval every change so silent regressions can’t reach users.
  • Defend the whole design in a system-design interview as a chain of guarantees, not a box-and-arrow diagram.

Up next track: Production RAG — ground the model in your own data, with citations and evals.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.