Lesson 4 of 6 · 50 min

Building a reliable LLM client

The provider will time out, rate-limit, 500, and occasionally go down. The four-layer reliability stack — timeouts, classified retries with jitter, fallback, circuit breaker — plus idempotency, hedging, the gateway, and the interview round on resilient distributed clients.

The provider is not your dependency — it’s your weather

A model API will time out, return 429s after a hot launch, throw 5xx, and occasionally have a full outage that looks identical to a rate limit from the outside. A naive client turns each of those into a user-facing error or a doubled bill. The fix is a small, ordered stack of independently-testable layers — and it’s the same distributed-systems resilience problem interviewers have probed for a decade, now applied to an upstream that is unusually slow, unusually expensive per call, and unusually prone to silent change.
What makes an LLM call different from a normal RPC, and why the standard playbook needs adapting: calls are seconds, not milliseconds (so a hung call ties up a worker far longer), each call costs real money (so a blind retry literally doubles a bill), failures are often partial (a stream that dies at token 800 of 1,000), and an outage can look identical to a rate limit from the outside (both surface as errors with no body you can trust). Every layer below is the classic resilience pattern, tuned for those four facts.

The four-layer reliability stack (in order)

  1. 011. Per-call timeout — bound how long any single attempt can hang (connect + read).
  2. 022. Retry — only on retryable errors, with capped exponential backoff + jitter, honouring Retry-After.
  3. 033. Fallback — a chain to a different model / region / provider when retries are exhausted.
  4. 044. Circuit breaker — short-circuit a known-bad endpoint so it stops eating retries and fallbacks.
The order matters: the circuit breaker sits above retries and below fallbacks — it stops you retrying a dead endpoint and frees the fallback chain to reach a healthy one. AWS’s Builders Library and the Maxim production guide both codify this exact stack.
code
1REQUEST FLOW through the stack (one logical call)23  request4    -> [circuit breaker]  open?  --yes-->  skip endpoint A, go straight to fallback5         |  closed6         v7       [timeout-bounded attempt]  --ok-->  return8         |  retryable error9         v10       [retry w/ backoff+jitter]  attempts left?  --yes-->  wait, re-attempt11         |  exhausted (and record failure -> may trip the breaker)12         v13       [fallback chain]  cheaper model / other region / other provider14         |  all exhausted15         v16       surface a clean error to the caller (never a raw 500)1718  Each layer is independently testable: fault-inject one, assert the next engages.

Timeouts: connect + per-request, sized from p99.9

Set both a connect timeout (short) and a read timeout, and size the read timeout near the provider’s p99.9 latency — that gives ~0.1% false-timeout rate while it’s healthy. Separate two budgets: the end-to-end budget the user feels (e.g. an 8s TTFT target, as Intercom Fin tracks) from the per-call budget on the HTTP client (35–50s including retry round-trips). LiteLLM defaults to a 30s per-call timeout.
A timeout subtlety that separates seniors: with streaming, the right knob is usually a time-to-first-token timeout plus an idle/inter-token timeout, not one fixed read timeout for the whole response. A correct long answer can legitimately take 30s to finish streaming; what you actually want to detect is “the stream stalled” (no token for N seconds) and “the first token never came” (TTFT blew the budget) — those are the real failure signals, and a single read timeout conflates them with a slow-but-healthy completion. And always pair a timeout with server-side cancellation: when you abandon a call, cancel it upstream too, or you keep paying for tokens nobody will read.

Retries: classify, back off, jitter, honour Retry-After

Classify first: retryable = 429, 500, 502, 503, 504; non-retryable = 400 (including context-length overflow — it won’t fix itself), 401, 403, 404. Use capped exponential backoff with jitter — the jitter is what prevents a thundering herd of clients retrying in lockstep onto a recovering provider. Cap at ~3 attempts (LiteLLM’s default); beyond 5 you mostly just double the bill. And always honour the Retry-After header on a 429.
Why jitter specifically — this is a favourite interview probe. Without it, every client that failed at the same instant computes the same backoff (1s, 2s, 4s…) and re-hits the recovering provider in synchronised waves, re-tripping it indefinitely; this is the “thundering herd.” AWS’s Builders Library shows full jitter — sleep a random amount in [0, backoff] — flattens those waves into a smooth load and dramatically cuts total contention versus plain exponential backoff. The counter-intuitive lesson interviewers want: adding randomness makes a distributed system more stable, not less.
python
1import httpx, random2from tenacity import retry, stop_after_attempt, wait_exponential_jitter, retry_if_exception34RETRYABLE = {429, 500, 502, 503, 504}56def is_retryable(e) -> bool:7    return isinstance(e, httpx.HTTPStatusError) and e.response.status_code in RETRYABLE89@retry(stop=stop_after_attempt(3),10       wait=wait_exponential_jitter(initial=1, max=30),   # backoff + JITTER11       retry=retry_if_exception(is_retryable), reraise=True)12def call_llm(client: httpx.Client, payload: dict):13    r = client.post("/v1/chat/completions", json=payload,14                    timeout=httpx.Timeout(connect=3.0, read=45.0))  # connect + read15    if r.status_code == 429 and "retry-after" in r.headers:16        raise httpx.HTTPStatusError("rate limited", request=r.request, response=r)  # honour header upstream17    r.raise_for_status()18    return r.json()
The streaming gotcha that bites teams: a call that already streamed 800 tokens and then dropped is not safe to blindly retry — retrying re-runs the whole generation (you pay for 800 + a fresh full completion, and the user may see duplicated text). Treat a mid-stream failure as a distinct case: either fail it cleanly, or only retry if you can resume, and never count partial output as a free do-over. This is the LLM-specific twist on idempotency that a generic retry library won’t handle for you.

Rate limits, fallback, circuit breakers

Two budgets must converge: the provider’s TPM/RPM (signalled by 429 + Retry-After) and a client-side token-bucket that refuses an over-budget request before it leaves your process — this is what actually prevents a rate-limit storm. Fallback is a decision tree: same provider / cheaper model (cheapest) → different region/tier → different provider (expect quality drift and prompt-portability work). The circuit breaker opens after N consecutive failures (typ. 5–10 in 60s), stays open 30–60s, then allows one half-open probe. Intercom Fin watches provider outages and a distinct “wasted tokens” metric via SLOs.
The circuit breaker is a small state machine worth being able to draw: closed (requests flow; count failures over a window) → trips to open after the threshold (requests fail fast for a cooldown, sparing the dead endpoint and your latency) → after the cooldown, half-open (allow a single probe; success closes it, failure re-opens). The senior subtlety is the fallback quality gap: failing over to a cheaper or different-provider model keeps you up, but the answer may be worse and your strict-mode prompt may not port cleanly — so a real fallback path needs its own (lighter) eval and its prompt validated against that model, or you trade an outage for a silent quality regression that’s harder to notice.
code
1CIRCUIT BREAKER state machine23           failures >= threshold (e.g. 5 in 60s)4   CLOSED  ----------------------------------------->  OPEN5   (flow,    <----------------------------------       (fail fast,6    count)        probe succeeds                         cooldown 30-60s)7      ^                                                    |8      |                                                    | cooldown elapsed9      |              probe fails                           v10      +------------------------------------------------  HALF-OPEN11                                                          (allow ONE probe)1213  Open state is the point: it STOPS you spending retries+timeouts on a known-dead14  endpoint, which is exactly the spend that turns one provider blip into your outage.

Hedging: trading a little cost for tail latency

A technique senior candidates score points with: hedged requests (Dean & Barroso, “The Tail at Scale”). If a call hasn’t returned a first token by, say, the p95 latency, fire a second request and take whichever responds first, cancelling the loser. This collapses the long tail — the occasional 8s straggler becomes ~p95 — at the cost of a small percentage of duplicate calls (you cap the extra spend by only hedging the slow tail, not every request). The catch for LLMs: hedging multiplies cost on a per-call-expensive upstream and is unsafe for non-idempotent tool calls, so reserve it for read-style, latency-critical paths and bound the hedge rate. Knowing when not to hedge is the senior signal.

Idempotency & the gateway

For write-side operations, use an idempotency key (OpenAI accepts Idempotency-Key; hash model+prompt) and a Redis dedup table at the gateway so two retrying workers collapse into one upstream call — no double charges, clean traces. And once more than one feature uses LLMs, stand up an AI gateway (LiteLLM self-hosted, Portkey, or Cloudflare AI Gateway) on day one: it centralises routing, retries, caching, spend caps, and observability, and it makes provider portability a config change rather than a rewrite.
The gateway is also where reliability becomes operable at org scale: one place to enforce per-team spend caps (so a runaway loop can’t burn the monthly budget in an hour), one place to emit the metrics that matter (success rate, retry rate, fallback rate, p50/p95/p99 TTFT, and a distinct wasted-tokens counter for abandoned/cancelled calls), and one place to flip a provider when one degrades. Centralising it means every feature inherits the hardened client instead of each team re-implementing — badly — its own retries and breaker.

Case studies & scale: reliability as an SLO discipline

Intercom Fin treats this as SLO engineering: an end-to-end TTFT target (around 8s) the user actually feels, separate from per-call HTTP budgets, plus monitoring on provider outages and a distinct “wasted tokens” metric for completions that get generated but never read (abandoned requests, fired-and-forgotten hedges). Klarna is the cautionary tale on the people side: an AI assistant that handled roughly two-thirds of chats and was credited with a ~$40M profit signal, then walked back toward human agents when quality slipped — the engineering lesson is to ship the regression detector and a graceful dial-back-to-human path before the savings headline. And the cross-cutting pattern from teams running this at volume: the failure that actually pages you is rarely a clean 500 — it’s a partial degradation (elevated latency, a creeping 429 rate, a fallback model quietly serving worse answers) that only structured metrics and a circuit breaker make visible before users complain.
Treat the provider as weather, not a dependency: you don’t prevent the storm, you build the client that stays up through it — bounded, jittered, breaker-guarded, and observable — and you measure the wasted tokens it costs you to do so.

Interview prep

This lesson maps directly onto a classic systems-design round: “design a resilient client for a flaky, slow, expensive upstream.” Interviewers test whether you reach for the ordered stack (not just “add retries”), whether you can explain why each layer exists and why the order matters, and whether you know the LLM-specific twists (cost-doubling retries, partial streams, fallback quality drift, wasted tokens). Lead with the mechanism, name the failure it prevents, and add a number where you have one.
  1. 01“Design a reliable LLM client.” → ordered stack: per-call timeout (connect+read) → classified retry (backoff+jitter, honour Retry-After) → fallback chain → circuit breaker; plus client-side rate limiter, idempotency, gateway.
  2. 02“Why does the order matter?” → breaker above retries (stop retrying a corpse), below fallback (free the chain to reach a healthy endpoint); retry for blips, breaker for sustained failure, fallback for ‘this one is down.’
  3. 03“Why jitter, not just exponential backoff?” → without it, clients that failed together retry in lockstep and re-trip the recovering provider (thundering herd); full jitter spreads load — randomness makes the system more stable.
  4. 04“Which errors do you retry?” → retry 429/500/502/503/504; never retry 400 (incl. context-length overflow), 401, 403, 404 — they fail identically next time.
  5. 05“How do you size a timeout?” → near the provider’s p99.9 for ~0.1% false timeouts; for streaming use a TTFT timeout + idle/inter-token timeout, not one whole-response read timeout; cancel server-side on abandon.
  6. 06“How do you stop a rate-limit storm?” → a client-side token bucket that refuses over-budget calls before they leave the process, plus jitter and a breaker — not more retries.
  7. 07“What’s a circuit breaker and its states?” → closed (count failures) → open (fail fast, cooldown) → half-open (one probe) → close on success; it stops spending retries/timeouts on a dead endpoint.
  8. 08“What’s special about retrying an LLM call?” → cost doubles on every retry, and a mid-stream failure already burned tokens — don’t blindly re-run; classify and, for streams, fail clean or resume.
  9. 09“What is request hedging and when do you avoid it?” → fire a second call past p95 and take the winner to cut tail latency; avoid on expensive per-call upstreams and non-idempotent tool calls, and bound the hedge rate.
  10. 10“Where does a gateway fit?” → one hardened client for all features: routing, retries, caching, spend caps, idempotency dedup, and the metrics (success/retry/fallback rate, p99 TTFT, wasted tokens) — provider swap becomes config.
Push it deeper. Expect: “Your fallback provider is up but answers are worse — how do you catch that?” (fallback paths need their own light eval and a prompt validated for that model; alert on a quality metric, not just availability). “Two workers retry the same write and you get a double charge — fix?” (idempotency key = hash of model+prompt+request id, deduped at the gateway so concurrent retries collapse to one upstream call). “How do you tell a provider outage from a rate-limit?” (you often can’t from one response — use the breaker’s failure window and Retry-After presence; treat sustained errors as outage and fail over). “What do you alert on?” (success rate, retry rate, fallback rate, p99 TTFT, and wasted-tokens — partial degradation, not just 500s, is what pages you).
articleTimeouts, retries and backoff with jitter (the canonical reference)Amazon Builders’ Libraryrepoopenai-cookbook — handling rate limits, retries, backoffopenai/openai-cookbookrepoLiteLLM — gateway: routing, retries, fallbacks, budgetsBerriAI/litellmpaperThe Tail at Scale (hedged requests; taming p99 latency)Dean & Barroso (CACM)

Checkpoint

Right after a launch, the provider starts returning 429s; your clients retry immediately, and the problem gets worse, not better. What’s the fix?

AIncrease the retry count so requests eventually succeedBAdd capped backoff + jitter, a client-side rate limiter that refuses over-budget calls, and a circuit breakerCSwitch every request to a second provider permanently
Sign up free to answer and see why

Checkpoint

Which of these should your client NOT retry?

AA 503 Service UnavailableBA 400 caused by the prompt exceeding the context windowCA 429 with a Retry-After header
Sign up free to answer and see why

Checkpoint

Your client streams responses. A call dies after streaming ~800 of an expected ~1,000 tokens. Your generic retry decorator re-fires the whole request. What’s the problem and the better design?

ATreat a mid-stream failure as a distinct case — fail it cleanly (or resume if supported), and don’t count the partial output as a free re-runBNo problem — retrying is always correct for transient failuresCLower the temperature so the stream doesn’t fail
Sign up free to answer and see why

Checkpoint

An interactive feature has a fat tail: p50 TTFT is 600ms but p99 is ~8s, and users hate the stragglers. Cost has some headroom. Strongest lever?

ARaise max_retries so slow calls get retriedBHedge the slow tail: if no first token by ~p95, fire a second request and take the winner — bounding the hedge rate so extra cost stays smallCIncrease the read timeout so slow calls have more time
Sign up free to answer and see why

Checkpoint

Your primary model is down so the breaker fails over to a cheaper second-provider model. Availability recovers, but a week later you learn answer quality quietly dropped during the outage. What was missing?

AA longer circuit-breaker cooldownBA higher retry count on the primaryCA fallback path with its own light eval and a prompt validated for that model, plus a quality metric alert — not just an availability check
Sign up free to answer and see why

Could you design, implement, and test the ordered reliability stack, explain why each layer and the order exist, handle the LLM-specific traps, and field the systems-design round?

Not yetMostlyConfident

Takeaways

  • Layer it in order: timeout → classified retry (backoff+jitter, honour Retry-After) → fallback → circuit breaker — each independently testable.
  • Size read timeouts near the provider’s p99.9; for streaming use a TTFT + idle timeout, and cancel server-side on abandon.
  • Jitter prevents the thundering herd; a client-side rate limiter — not more retries — is what stops a rate-limit storm.
  • LLM-specific traps: every retry doubles cost, partial streams already cost money, and a fallback model can silently serve worse answers.
  • Hedge the slow tail to tame p99 — bounded, and never on non-idempotent calls.
  • Use idempotency keys for write-side calls; run a gateway once you have >1 LLM feature, and alert on wasted tokens, not just 500s.

Next: making it fast and cheap — the cost & latency lever stack.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.