REST resources, status semantics, error envelopes, versioning, cursor pagination, rate limits, and async job patterns for LLM backends.
Lesson 1 · HTTP APIs
REST, status codes, versioning, pagination, error contracts
Pretty OpenAPI is not a production API
Interviewers and production both grade the contract under failure: wrong status codes, undocumentable error shapes, unbounded list endpoints, silent breaking changes, and versioning theater. AI features sit on the same rails — a streaming completion endpoint still needs auth, rate limits, idempotent side effects, and a client-safe error body. This lesson locks HTTP design you can defend in a design doc and a coding loop.
Backend for AI builders is not “call OpenAI from a route.” It is stable boundaries: resources, methods, status semantics, pagination, and errors that clients can branch on without reading your source. Mid answers list REST verbs. Senior answers name cacheability, partial failure, backwards compatibility, and what the client does on 429 vs 503.
Resources, not RPC-by-default
Model the noun first: job, run, document, embedding-job, webhook-delivery. Use verbs only when the domain truly is an action that is not a resource state change (POST /v1/search is fine; POST /doThing for every mutation is not). Prefer create resource → poll or stream status for long LLM work over holding an HTTP connection for minutes without a job id.
REST is a style, not a religion. gRPC is excellent service-to-service. GraphQL can reduce overfetch for product BFF shapes. The interview grade is whether you pick transport for the client (browser, mobile, worker, partner) and can explain caching, errors, and auth for that choice — not whether you recite Fielding from memory.
Status codes that clients can automate
Status codes are a machine contract. 2xx success families, 4xx client must change request, 5xx retry or page someone. Abuse collapses automation: returning 200 with {"error": true} forces every client to invent parsers; returning 500 for validation errors hides product bugs as outages.
text
1STATUS CHEAT SHEET (interview + prod)23200 OK full success body for GET/PUT/PATCH as designed4201 Created POST created; Location header when useful5202 Accepted async accepted; body has job/run id + poll URL6204 No Content success, empty body (DELETE often)7400 Bad Request malformed JSON / schema fail (not auth)8401 Unauthorized missing/invalid credentials9403 Forbidden authenticated but not allowed10404 Not Found resource missing OR hidden (authz privacy)11409 Conflict version/state conflict (optimistic lock)12422 Unprocessable semantic validation (domain rules)13429 Too Many rate limit; honor Retry-After14500 Internal bug; client may retry carefully15502/503/504 upstream/gateway/timeout; retry with backoff1617LLM feature tip: long generation → 202 + job resource, not 200 held open18unless you intentionally stream (SSE/WebSocket) with cancel + heartbeat.
Error contract — one shape forever
Pick one error envelope and freeze it. Clients need: stable code (machine), human message (safe to show), optional details (field errors), and request_id for support. Never leak stack traces or internal table names to external clients. Internally log the full cause under the same request_id.
json
1// Canonical error body (JSON APIs)2{3 "error": {4 "code": "rate_limited",5 "message": "Too many completion requests. Try again shortly.",6 "details": {"limit": 60, "window": "1m"},7 "request_id": "req_01J8Z..."8 }9}1011// Anti-patterns12// 200 { "success": false, "err": "something" } — unautomable13// 500 "pq: duplicate key value violates unique" — leak + wrong code14// Different shapes per route — multiplies client bugs
Versioning without theater
Version when you will break clients: remove fields, change semantics, tighten auth. Prefer additive evolution (new optional fields, new endpoints) over /v2 for every tweak. Common strategies: URL prefix (/v1/), header version, or date version (Stripe-style). Pick one primary strategy and document deprecation windows.
For internal-only services, a single unversioned surface can work if you own all clients and ship lockstep. For partner or mobile clients, assume year-long lag. Senior signal: name who cannot upgrade quickly before proposing a break.
Pagination — never unbounded lists
List endpoints without pagination are production incidents waiting for a power user. Prefer cursor pagination for large/mutable sets (stable under inserts). Offset/limit is simpler for admin UIs and small tables but drifts under concurrent writes and gets expensive on deep pages.
text
1GET /v1/runs?limit=50&cursor=eyJpZCI6Li4ufQ23{4 "data": [ /* up to 50 runs */ ],5 "pagination": {6 "next_cursor": "eyJpZCI6Li4ufQ",7 "has_more": true8 }9}1011# Cursor encodes (sort_key, id) so equal timestamps still advance.12# Never accept raw SQL offsets from the client.13# Default + max limit (e.g. default 20, max 100).
Idempotent methods and safe retries
GET/HEAD/PUT/DELETE are defined as idempotent in spirit (DELETE twice → still gone). POST is not — retries create duplicates unless you add Idempotency-Key (Stripe pattern) or make the resource identity client-supplied. Timeouts after POST are the classic double-charge / double-run bug for LLM jobs that bill tokens.
Rate limits and 429 honesty
Return 429 with Retry-After when the client should slow down. Separate product quotas (plan limits) from abuse protection (edge). Document limits in human units (requests/min, tokens/day). For LLM routes, budget on tokens or concurrent jobs — RPS alone underprices expensive completions.
Streaming and long work
Token streaming (SSE) is a UX choice with backend costs: sticky connections, cancel semantics, partial failure mid-stream. Alternative: 202 Accepted create run → client polls GET /runs/{id} or subscribes to events. Interviewers love hearing both paths and when you pick each (interactive chat vs batch classify).
If you stream: define heartbeat/comment frames so proxies do not idle-timeout; define cancel (client disconnect or explicit POST cancel) so you stop paying the provider; define what partial text means if the connection dies mid-completion (persist partial vs fail the run).
Content negotiation and headers that matter
Accept and Content-Type should be explicit (application/json). Use ETag/If-Match for optimistic concurrency on mutable resources. Correlation: generate request_id at the edge and echo it on every response. Deprecation: Sunset/Deprecation headers when you retire fields. These small headers separate production APIs from demo routes.
Multi-tenant API hygiene
Never take tenant_id only from the body if auth already knows the tenant — prefer deriving it from the credential. Path params like /tenants/{tid}/… still need an authz check that the caller belongs there. Cross-tenant resource ids in URLs are the classic IDOR footgun you will see in L3, but the API design starts here: stable resource paths and consistent 404 vs 403 policy.
If the client cannot tell success, retryable failure, and permanent failure apart from your status + body alone, the API is incomplete.
Interview answers — HTTP & API design
01Q: REST vs RPC? REST when resources + cache + uniform errors matter to many clients; RPC/gRPC when internal latency and typed contracts dominate. Say who the client is before picking.
02Q: 401 vs 403? 401 means we do not know who you are (or credentials failed). 403 means we know who you are and you still may not do this. Do not use 401 to hide existence if that confuses login flows.
03Q: Why cursor pagination? Stable under inserts, O(1)-ish page advance with the right index, avoids deep OFFSET cost. Offset is fine for small stable admin lists.
04Q: How do you version APIs? Additive first; URL or header version for breaks; deprecation window; never silently change field meaning under the same version.
05Q: POST timeout — what happens? Assume at-least-once client retry. Design idempotency keys or client-generated ids so a second POST does not double-bill or double-enqueue.
06Q: When 202 vs streaming? 202 + job when work is long or batch; stream when UX needs tokens as they arrive and you can cancel/cleanup connections.
07Q: Error body essentials? Stable code, safe message, request_id, optional field details. Same envelope everywhere.
08Q: What breaks mobile clients? Removing fields, changing nullability, renumbering enums, requiring new headers without versioning.
09Q: Rate limit design? Token bucket at gateway; 429 + Retry-After; separate expensive LLM routes; identity key = user/api_key not only IP.
10Q: HATEOAS? Rarely full hypermedia in product APIs; still return explicit next links/cursors so clients do not invent URL rules.
11Q: GraphQL pitfalls? Authz per field, N+1 resolvers, unbounded query cost — need depth/complexity limits and DataLoader-style batching.
12Q: Public vs internal API? Public needs stricter versioning, abuse controls, and docs; internal can be chattier but still needs stable errors and auth.
A client POSTs /v1/embedding-jobs, the server enqueues work, then the client times out before reading the body. The client retries the same POST. What design prevents duplicate jobs?
AReturn 500 on the first timeout so the client knows to stop retrying permanently.BRequire an Idempotency-Key (or client-generated job id) so the second POST returns the original job instead of creating another.CSwitch the endpoint to GET so HTTP caching makes retries free.
Your list endpoint returns all matching rows with no limit. A tenant with 2M rows opens the admin UI. What is the production failure mode?
AOnly a UX issue — databases handle large SELECTs fine if you have indexes.BUnbounded response size and query cost — require cursor/limit with a max page size and stable sort keys.CFix by returning HTTP 204 always so payloads stay small.
A validation failure on an LLM run create currently returns HTTP 500 with a SQL error string. What should it return instead?
AHTTP 422 or 400 with a stable error code, safe message, and request_id — never raw SQL in external bodies.BHTTP 200 with {ok:false} so existing clients that only check status keep working.CHTTP 503 so the client retries until the prompt becomes valid.
You need to change a response field from string to object. Mobile clients lag 6 months. Best approach?
AShip the break under /v1 immediately; mobile will crash and users will update.BAdd a new field or /v2, keep old field until deprecation window ends, document the sunset, then remove.COnly document the change in Slack; status codes stay the same so it is not breaking.
When is 202 Accepted the better default than holding a connection for a multi-minute LLM batch?
AAlways — streaming is never appropriate for user-facing products.BWhen work can outlive the client connection and you expose a job/run resource to poll or subscribe to.COnly when the database is down and you cannot write results yet.