Lesson 8 of 8 · 55 min
Capstone: design a reliable LLM feature backend
End-to-end mock: Diff Explain SaaS — API, data model, authz, queues, idempotent billing, caching, SLOs, and a 45-minute self-scored rubric.
Lesson 8 · Capstone
Design a reliable LLM feature backend end-to-end
Wire every lesson into one system
Capstone prompt (use in mocks)
1PRODUCT: Diff Explain (SaaS)2Users paste/upload a unified diff (≤ 200KB). System returns:3 - summary of risk4 - bullet explanation of changes5 - optional test suggestions6Constraints:7 - multi-tenant; projects; RBAC (admin/member/viewer)8 - p99 interactive path for small diffs OR async for large9 - bill tokens to the tenant; hard budget caps10 - CI can call via API key + optional webhook on completion11 - provider may 429/5xx; must not double-bill on retries12 - on-call must detect lag, error spikes, cost spikes1314Deliver in 45 minutes whiteboard:15 1. API 2. Data model 3. Auth 4. Async plan16 5. Idempotency 6. Perf 7. Observability 8. RisksReference architecture (senior skeleton)
1CLIENT2 │ POST /v1/diff-explains Idempotency-Key3 ▼4API GATEWAY (authn, rate limit, request_id)5 │6 ▼7CONTROL API8 ├─ AuthZ project:write9 ├─ TX: insert run queued + outbox + idem row10 └─ 202 { run_id, status_url }11 or SSE upgrade for small/interactive12 │13 ▼14QUEUE ── worker ── provider LLM15 │ │16 │ ├─ object storage (explanation.md)17 │ ├─ DB status + token ledger (unique run_id)18 │ └─ webhook dispatcher (event id dedupe)19 ▼20GET /v1/diff-explains/{id} (poll) + metrics/traces everywhereKey idea
API contract (minimum)
Data model (minimum)
1-- Status machine (enforce in UPDATE WHERE)2queued → running → succeeded3 → failed4queued|running → cancelled56usage_ledger (7 run_id PK REFERENCES diff_runs(id),8 input_tokens INT NOT NULL,9 output_tokens INT NOT NULL,10 usd_micros BIGINT NOT NULL11) -- one ledger row per run: natural idempotencyCommon mistake
“We will store status only in the queue message attributes.”
AuthN/Z on the critical path
Async, retries, DLQ
Idempotency map
Key idea
Caching & perf
Observability & SLOs
Common mistake
“If the model answer is high quality, backend reliability is optional for MVP.”
Worked mock: 45-minute agenda
10–5m Clarify: size limits, sync vs async UX, tenants, CI, budgets25–12m API + error/status codes + idempotency header312–20m Data model + status machine + ledger420–28m AuthZ + API keys + webhook signing528–36m Workers, retries, DLQ, cancel/spend stop636–42m SLOs, dashboards, failure modes (provider 429, poison)742–45m Tradeoffs & what you would build in week 1 vs laterSelf-score rubric (100 pts)
Failure mode drill (say these out loud)
1FAILURE MITIGATION SIGNAL2------------------- ---------------------------- --------------------3Provider 429 backoff, queue, message UX provider_error_rate4Worker crash visibility redelivery lag + DLQ5Redis down degrade flags/session path redis up + api p996PG failover pool retry, short errors db errors + p997Stolen API key revoke, rotate, audit auth anomalies8Webhook 500s retry + DLQ + inbox UI delivery failures9Budget exceeded hard stop new runs spend + 402/42910Poison payload validate + DLQ crash fingerprintWeek-1 vs later roadmap
What interviewers dock points for
Interview answers — capstone synthesis
- 01Q: Sync or async? Small interactive SSE optional; default 202+job for reliability and CI; never block load balancers for minutes without a plan.
- 02Q: System of record? Postgres run status + ledger; object storage for large text; queue is transport.
- 03Q: Double bill risk? Idempotent create + unique ledger on run_id + careful provider retries.
- 04Q: Cross-tenant leak? tenant/project checks on every read; tests with foreign UUIDs; no shared cache keys.
- 05Q: Provider outage? Retry/backoff, circuit break, degrade messaging, error budget, status page honesty.
- 06Q: Poison diff? Validation, size caps, DLQ, do not infinite crash loop workers.
- 07Q: Week-1 cut? Single region, one worker pool, Stripe-like idempotency, basic RED metrics — postpone multi-region and fancy caches.
- 08Q: Webhook reliability? Sign, retry with backoff, dedupe event ids, show delivery logs.
- 09Q: Cost control? Per-tenant budgets, kill switch, anomaly alerts, model allowlists.
- 10Q: What fails first at 10×? Provider rate limits, queue lag, DB list queries without indexes — name mitigations.
- 11Q: Cancel semantics? Cooperative cancel between steps; mark cancelled; avoid orphan spend when possible.
- 12Q: Security review? Secrets, authz tests, SSRF if tools fetch, prompt log redaction, key hashing.
Checkpoint
In Diff Explain, where should “tokens billed to tenant” be recorded to prevent double billing under worker redelivery?
Checkpoint
CI posts the same Diff Explain request twice after a network timeout. What design makes this safe?
Checkpoint
Viewer role calls POST /v1/diff-explains on a project. Correct response?
Checkpoint
Which SLO set best matches this product’s async nature?
Checkpoint
Week-1 MVP cut: what do you postpone without abandoning reliability rails?
Can you whiteboard Diff Explain covering API, data, authz, async, idempotency, perf, and SLOs in 45 minutes?
Takeaways
- LLM features inherit all backend rails — they do not replace them.
- Status DB + queue transport + unique ledger is the reliability spine.
- Idempotency and object-level authz are non-negotiable MVP pieces.
- SLOs must include async lag and cost, not only API CPU.
Track complete. Re-run the capstone mock weekly and drill weak lessons (auth, idempotency, obs) before loops.
Sources
Free to read · better with Enzo
Learn it with Enzo
Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.