Lesson 8 of 8 · 55 min

Capstone: design a reliable LLM feature backend

End-to-end mock: Diff Explain SaaS — API, data model, authz, queues, idempotent billing, caching, SLOs, and a 45-minute self-scored rubric.

Lesson 8 · Capstone

Design a reliable LLM feature backend end-to-end

Wire every lesson into one system

Prompt: design the backend for “Explain this diff” — authenticated users submit a git diff, get a streaming or async explanation, billed per run, with webhooks for CI. You will produce API shapes, data model, authz, queue flow, idempotency, caching, and observability. This is the mock loop interviewers run for AI platform / FDE / backend roles.
Grade yourself on the track rubric: contracts under failure, invariants in the DB, object-level authz, at-least-once with effect-once, measured hot paths, and on-call signals. Pretty model diagrams without these are junior.

Capstone prompt (use in mocks)

text
1PRODUCT: Diff Explain (SaaS)2Users paste/upload a unified diff (≤ 200KB). System returns:3  - summary of risk4  - bullet explanation of changes5  - optional test suggestions6Constraints:7  - multi-tenant; projects; RBAC (admin/member/viewer)8  - p99 interactive path for small diffs OR async for large9  - bill tokens to the tenant; hard budget caps10  - CI can call via API key + optional webhook on completion11  - provider may 429/5xx; must not double-bill on retries12  - on-call must detect lag, error spikes, cost spikes1314Deliver in 45 minutes whiteboard:15  1. API  2. Data model  3. Auth  4. Async plan16  5. Idempotency  6. Perf  7. Observability  8. Risks

Reference architecture (senior skeleton)

Edge: TLS, WAF, rate limit. API: auth session or API key, validate size, create run (idempotent), enqueue, 202 or start SSE for small mode. Workers: lease job, conditional status, call provider with timeout/cancel, store artifact in object storage, ledger tokens, fire webhooks idempotently. Control data in Postgres; hot session/flags in Redis; queue SQS/equivalent.
text
1CLIENT2  │  POST /v1/diff-explains  Idempotency-Key3  ▼4API GATEWAY (authn, rate limit, request_id)5  │6  ▼7CONTROL API8  ├─ AuthZ project:write9  ├─ TX: insert run queued + outbox + idem row10  └─ 202 { run_id, status_url }11        or SSE upgrade for small/interactive12  │13  ▼14QUEUE ── worker ── provider LLM15  │         │16  │         ├─ object storage (explanation.md)17  │         ├─ DB status + token ledger (unique run_id)18  │         └─ webhook dispatcher (event id dedupe)19  ▼20GET /v1/diff-explains/{id}  (poll) + metrics/traces everywhere

API contract (minimum)

POST create with Idempotency-Key; GET run; GET artifact; POST cancel; webhook config on project. Errors: 401/403/404/409/413/422/429/5xx with one envelope. Pagination for list runs with cursor. Version /v1.

Data model (minimum)

tenants, users, memberships(role), projects, api_keys(hash), diff_runs(status, model, token counts, error_code, version), idempotency_keys, outbox, webhook_endpoints, webhook_deliveries(event_id unique), usage_ledger(unique run_id). Artifacts by pointer.
text
1-- Status machine (enforce in UPDATE WHERE)2queued → running → succeeded3                 → failed4queued|running → cancelled56usage_ledger (7  run_id PK REFERENCES diff_runs(id),8  input_tokens INT NOT NULL,9  output_tokens INT NOT NULL,10  usd_micros BIGINT NOT NULL11)  -- one ledger row per run: natural idempotency

AuthN/Z on the critical path

Session for web; hashed API keys for CI. Every get/cancel checks project membership. Viewer cannot create. Webhooks signed outbound with secret. Provider keys only on workers via secret manager.

Async, retries, DLQ

Default async 202 for reliability. Retry provider 429/5xx with backoff and budget. Terminal model errors → failed without infinite retry. DLQ + alert. Cancel sets flag workers honor between steps to stop spend.

Idempotency map

Create-run key from client. Worker processes run_id once for ledger insert. Webhook deliveries unique on event_id. Provider calls: include run_id in metadata; if vendor supports idempotency keys, pass them.

Caching & perf

Cache project config and feature flags. Do not cache explanations under a key that drops tenant_id. Keyset list runs. N+1 banned on list endpoints. Size limit 200KB with 413. Optional: cache identical diff hash results inside a tenant only with explicit policy.

Observability & SLOs

SLIs: create success, time-to-running, time-to-terminal, provider error rate, token spend/hour/tenant, DLQ depth. Trace API→queue→worker→provider. Page on burn + lag + spend anomaly. Runbooks: provider outage, poison payload, budget lock.

Worked mock: 45-minute agenda

text
10–5m   Clarify: size limits, sync vs async UX, tenants, CI, budgets25–12m  API + error/status codes + idempotency header312–20m Data model + status machine + ledger420–28m AuthZ + API keys + webhook signing528–36m Workers, retries, DLQ, cancel/spend stop636–42m SLOs, dashboards, failure modes (provider 429, poison)742–45m Tradeoffs & what you would build in week 1 vs later

Self-score rubric (100 pts)

API contract 15 · Data/invariants 15 · Authz 15 · Async ops 15 · Idempotency/billing 15 · Perf 10 · Observability 10 · Communication 5. Senior bar ~80+ with explicit tradeoffs. Mid bar lists components without failure modes.

Failure mode drill (say these out loud)

Provider 429 for 15 minutes; worker deploy kills mid-job; Redis down; Postgres primary failover; malicious 200KB-looking zip bomb if you ever accept binaries; stolen API key; webhook endpoint returning 500 for a day; double-click submit; tenant budget exceeded mid-run. For each: user visible effect, mitigation, metric that pages.
text
1FAILURE              MITIGATION                     SIGNAL2-------------------  ----------------------------  --------------------3Provider 429         backoff, queue, message UX    provider_error_rate4Worker crash         visibility redelivery        lag + DLQ5Redis down           degrade flags/session path   redis up + api p996PG failover          pool retry, short errors     db errors + p997Stolen API key       revoke, rotate, audit        auth anomalies8Webhook 500s         retry + DLQ + inbox UI       delivery failures9Budget exceeded      hard stop new runs           spend + 402/42910Poison payload       validate + DLQ              crash fingerprint

Week-1 vs later roadmap

Week 1: single region, one worker queue, Postgres+Redis+SQS-class, idempotent create, ledger unique, basic RED+lag+spend alerts, authz tests. Later: multi-region, fine-grained priority queues, advanced eval pipelines, customer-facing delivery logs UI, automatic poison replay. Interviews reward honest sequencing.

What interviewers dock points for

No authz on GET by id; fire-and-forget threads without durability; “exactly-once Kafka” as a substitute for ledger design; caching explanations without tenant in the key; paging only on CPU; ignoring cancel/spend; no size limits on diffs; hand-waving webhooks. If your design hits three of these, rewrite before the critique round.

Interview answers — capstone synthesis

  1. 01Q: Sync or async? Small interactive SSE optional; default 202+job for reliability and CI; never block load balancers for minutes without a plan.
  2. 02Q: System of record? Postgres run status + ledger; object storage for large text; queue is transport.
  3. 03Q: Double bill risk? Idempotent create + unique ledger on run_id + careful provider retries.
  4. 04Q: Cross-tenant leak? tenant/project checks on every read; tests with foreign UUIDs; no shared cache keys.
  5. 05Q: Provider outage? Retry/backoff, circuit break, degrade messaging, error budget, status page honesty.
  6. 06Q: Poison diff? Validation, size caps, DLQ, do not infinite crash loop workers.
  7. 07Q: Week-1 cut? Single region, one worker pool, Stripe-like idempotency, basic RED metrics — postpone multi-region and fancy caches.
  8. 08Q: Webhook reliability? Sign, retry with backoff, dedupe event ids, show delivery logs.
  9. 09Q: Cost control? Per-tenant budgets, kill switch, anomaly alerts, model allowlists.
  10. 10Q: What fails first at 10×? Provider rate limits, queue lag, DB list queries without indexes — name mitigations.
  11. 11Q: Cancel semantics? Cooperative cancel between steps; mark cancelled; avoid orphan spend when possible.
  12. 12Q: Security review? Secrets, authz tests, SSRF if tools fetch, prompt log redaction, key hashing.
articleStripe IdempotencyStripedocsGoogle SRE — MonitoringGoogle SREdocsOWASP API Security Top 10OWASParticleTransactional Outboxmicroservices.io

Checkpoint

In Diff Explain, where should “tokens billed to tenant” be recorded to prevent double billing under worker redelivery?

AOnly in a metric counter incremented on every provider HTTP attempt.BIn a usage_ledger row uniquely keyed by run_id (same transaction as success transition when possible).COnly in the client’s browser localStorage for transparency.
Sign up free to answer and see why

Checkpoint

CI posts the same Diff Explain request twice after a network timeout. What design makes this safe?

ADisable all CI retries in documentation and hope vendors comply.BRequire Idempotency-Key on create; second POST returns the original run without enqueueing a duplicate.CUse GET with the diff in query string so HTTP caches dedupe automatically.
Sign up free to answer and see why

Checkpoint

Viewer role calls POST /v1/diff-explains on a project. Correct response?

A403 Forbidden after authenticating the user and failing project permission check.B201 Created because authentication succeeded and viewers should always write.C500 with a stack trace naming the permission table.
Sign up free to answer and see why

Checkpoint

Which SLO set best matches this product’s async nature?

AOnly CPU < 50% on API hosts, checked daily by hand.BCreate-run API availability/latency, time-to-running lag, terminal success rate, spend anomaly guards.COnly model BLEU score on a research dataset.
Sign up free to answer and see why

Checkpoint

Week-1 MVP cut: what do you postpone without abandoning reliability rails?

APostpone multi-region active-active and sophisticated CDN; keep idempotency, authz checks, outbox/queue, ledger uniqueness, basic RED/lag alerts.BPostpone authz and idempotency until after launch traffic arrives.CPostpone the database; store runs only in ephemeral worker memory.
Sign up free to answer and see why

Can you whiteboard Diff Explain covering API, data, authz, async, idempotency, perf, and SLOs in 45 minutes?

New to itGetting thereConfident

Takeaways

  • LLM features inherit all backend rails — they do not replace them.
  • Status DB + queue transport + unique ledger is the reliability spine.
  • Idempotency and object-level authz are non-negotiable MVP pieces.
  • SLOs must include async lag and cost, not only API CPU.

Track complete. Re-run the capstone mock weekly and drill weak lessons (auth, idempotency, obs) before loops.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.