Lesson 6 of 6 · 52 min
Capstone: a summarize-and-extract microservice
Assemble the whole track — prompt contract + structured output + reliable client + caching + tests + observability — into a service that turns a document into a structured summary you can trust in production, and that you can defend end-to-end in a system-design interview.
Capstone
A summarize-and-extract microservice
Where the five lessons become one service
The architecture
1REQUEST PATH (each stage = one lesson)23 doc in4 -> assemble prompt L2 stable cached system contract + JIT context5 -> semantic cache lookup L5 hit? return cached structured result, skip the model6 -> reliable client call L4 timeout -> retry(jitter) -> fallback -> breaker7 -> strict structured out L3 constrained decoding into Pydantic (no parse step)8 -> business-rule validate L3 validators; targeted re-ask on semantic failure only9 -> write caches + span L5 provider prompt cache always; log tokens/cost/latency10 structured result out1112 Cross-cutting: pinned model+prompt VERSION (L1/L2) + golden-set eval gate in CI1from pydantic import BaseModel2import instructor, time3from openai import OpenAI45class DocSummary(BaseModel):6 summary: str7 action_items: list[str]8 owner: str | None910client = instructor.from_openai(OpenAI(timeout=45.0, max_retries=3)) # L4 reliability1112SYSTEM = "Summarize the document and extract action items + owner. Use only the document." # L2 stable prefix1314def summarize(doc: str) -> tuple[DocSummary, dict]:15 t0 = time.time()16 out = client.chat.completions.create(17 model="gpt-4o-mini", # L5: cheap tier; escalate only the hard tail18 response_model=DocSummary, # L3: strict structured output19 max_retries=2, # business-rule retries only20 messages=[{"role": "system", "content": SYSTEM},21 {"role": "user", "content": doc}],22 )23 span = {"latency_ms": int((time.time()-t0)*1000)} # L5: log cost/latency per request24 return out, spanHandling a document bigger than the window
1def summarize_large(doc: str, chunk) -> DocSummary:2 parts = [summarize(c)[0].summary for c in chunk(doc)] # MAP: per-chunk (parallel)3 while n_tokens("\n".join(parts)) > WINDOW_BUDGET:4 parts = [summarize("\n".join(g))[0].summary # REDUCE: fold partials5 for g in groups_of(parts, 5)]6 return summarize("\n".join(parts))[0] # final pass over the foldTesting a probabilistic service: three concentric rings
1THREE TEST RINGS (cheapest + most reliable first)23 Ring What it checks Cost Run on4 ---- ------------------------------------ ---- ---------------------5 1 det schema valid, JSON valid, regex ids, free every request + CI6 exact-match closed fields (assert these always)7 2 sem embedding-similarity floor on summary, cheap CI + sampled prod8 refusal-string on safety prompts9 3 jdg stronger model scores vs a rubric, $$ CI gate + small sample10 calibrated to a HUMAN golden set (sampled to bound cost)1112 Pitfall: an UNcalibrated LLM judge drifts and is biased (length, position).13 Anchor it to ~50 human labels and re-check agreement, or it lies to you.model and prompt_version, score every change, and you convert “users tell you it got worse” into “CI blocks the PR.”Key idea
Observability: the span that explains every request
prompt_version, model, input/output token counts, computed cost, latency (TTFT + total), cache-hit flag, retry/fallback count, and whether validation repaired. With that, every production question becomes answerable: a cost cliff is a trend line on the cost field (L5), a quality dip bisects to a prompt_version (L2), a latency regression splits into prefill vs decode (L1), and an outage shows up as a fallback-rate spike (L4). Without it, you’re debugging a probabilistic black box from user complaints. This is why observability is a capstone requirement, not a nice-to-have.Common mistake
“It works in the demo — ship it.”
Case study: the same shape, at scale
A production LLM feature is not “a model that answers.” It’s a contract (schema), a shield (reliable client), a budget (cost/latency levers + span), and a gate (golden-set eval on pinned versions). Build those four and the model is just the easy part in the middle.
Interview prep
- 01“Design a document summarize-and-extract service.” → cached prompt contract → reliable client → strict structured output → validate/repair → cache + per-request span → golden-set eval gate on pinned versions.
- 02“How do you guarantee parseable output?” → strict structured outputs (constrained decoding into a Pydantic schema), not prompt-and-pray + try/except — removes the parse-failure class.
- 03“How do you test something nondeterministic?” → three rings: deterministic (schema/exact/regex), semantic (embedding floor), LLM-judge calibrated to ~50 human labels; ship ring 1 first, judge last and sampled.
- 04“What gates a deploy?” → a frozen golden set in CI scoring every prompt/model change against pinned model+prompt_version; block the merge on regression.
- 05“The doc is bigger than the window — what do you do?” → map-reduce for holistic summaries, chunk+retrieve for targeted extraction; not “bigger-context model” by reflex.
- 06“What do you log per request?” → prompt_version, model, in/out tokens, cost, TTFT+total latency, cache-hit, retry/fallback count — so every prod question is answerable.
- 07“Users say it got worse but you shipped nothing — what protected you?” → pinned versions + golden-set eval (Anthropic’s postmortem) turn a silent provider regression into a detectable, bisectable event.
- 08“How do you make it cheap?” → cache the stable prefix, route the easy majority to a cheap model, keep output short, batch any offline backfill — watch the per-request cost span.
- 09“Where does reliability live?” → the L4 stack (timeout → retry+jitter → fallback → breaker) plus idempotency for write-side calls; ideally behind a shared gateway.
- 10“What’s your rollout plan for a prompt change?” → canary on a traffic slice, compare on the golden set + live metrics, promote on no regression, rollback is a config flip.
Common mistake
The red flag that sinks candidates: stopping at “call the model and return the answer.”
request → model → response and stops has shown they’ve never operated one in production. The four missing pieces are exactly the senior signal: the output is an enforced schema, the upstream is wrapped in a reliability stack, cost/latency are engineered and observable (levers + a per-request span), and change is gated by a golden-set eval on pinned versions. Name those four unprompted and you’ve passed the round.Checkpoint
You’re writing the first tests for the service’s probabilistic output. Which ring do you ship first?
Checkpoint
Two weeks after launch, users report the service “got worse,” though you changed nothing. What would have protected you — and now tells you what happened?
Checkpoint
A user uploads a 400-page contract — well over the context window — and asks for a one-paragraph summary. Best approach for this service?
Checkpoint
Strict structured output is on, so JSON is always valid — yet the extracted owner field is wrong on ~30% of documents. Where is the problem and the fix?
Checkpoint
You add an LLM-as-judge as ring 3 and trust its scores to gate deploys. A month later it’s passing changes that users dislike. Most likely cause and fix?
Could you build, test, observe, and deploy-gate this microservice end-to-end, handle the bigger-than-window and accuracy cases, and defend the whole design in a system-design round?
You can now
- Assemble prompt contract + structured output + reliable client + caching + observability span into one service.
- Scale past the context window with map-reduce (summaries) or chunk+retrieve (targeted extraction).
- Test probabilistic output in three rings and gate deploys on a golden set with a calibrated judge.
- Pin versions and eval every change so silent regressions can’t reach users.
- Defend the whole design in a system-design interview as a chain of guarantees, not a box-and-arrow diagram.
Up next track: Production RAG — ground the model in your own data, with citations and evals.
Sources
Free to read · better with Enzo
Learn it with Enzo
Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.