Lesson 1 of 8 · 55 min

Building blocks every round reuses

Load balancers, caches, sharding, queues, IDs, CDN, search, WebSockets, observability — with reported/heuristic production numbers and a full request lifecycle.

Stop inventing boxes from scratch

Almost every classic system design prompt is a recombination of ~12 primitives — load balancer, cache, shard, replica, queue, ID generator, rate limiter, CDN, search index, object store, WebSocket gateway, and observability/cost. Interviewers grade whether you pick the right primitive for the pressure, not brand names. Memorize when each block is right, the failure it introduces, and a one-line tradeoff. Production numbers below are labeled reported (first-party eng blogs) or heuristic (interview ladder). The close annotates a full request lifecycle.
The round is shape recognition: read-heavy or write-heavy? Does the user need the latest write, or is eventual fine? Where is the fan-out? Where is the hot key? Name the pressure, name two options, pick a default, say what breaks. Do not memorize Kafka vs Pulsar trivia — reuse the toolkit. Seniors narrate why this box, not that one before they draw the next arrow.

Latency ladder (memorize once, reuse forever)

Jeff Dean’s classic latency table (widely circulated gist; treat as heuristic order-of-magnitude, not a lab measurement of your stack): L1 ~1 ns, branch ~3 ns, L2 ~5 ns, mutex ~17 ns, main memory ~100 ns, SSD random read ~100 µs, disk seek ~1 ms, same-DC RTT ~0.5 ms, inter-city ~5 ms, cross-ocean ~50 ms. The interview use: if your interactive p99 budget is 100 ms, you get roughly two cross-region RTTs total — so do not put a synchronous cross-ocean write on the hot path without saying so.
code
1LATENCY LADDER (heuristic — interview use, not a SLA)2  L1 cache hit            ~1 ns3  Branch mispredict       ~3 ns4  L2 hit                  ~5 ns5  Mutex lock/unlock      ~17 ns6  Main memory            ~100 ns7  SSD random read        ~100 us   (1e5 ns)8  Same datacenter RTT    ~0.5 ms9  SSD / NVMe sequential  higher throughput, still not free10  Disk seek              ~1–10 ms11  Same region cross-AZ   ~1–2 ms class12  Cross-city             ~5–20 ms13  Cross-ocean            ~50–150 ms1415DESIGN RULE: p99 budget ÷ RTT = how many remote hops you can afford.16If p99=100ms and hop=5ms, you do not get 30 sequential service calls.

Production building-block numbers (reported / heuristic)

code
1BLOCK                 WHAT IT IS                         NUMBER (label)2--------------------  ---------------------------------  ----------------------------------3DNS                   hierarchical lookup                cold multi-continent ~50–120 ms (heuristic);4                                                         Cloudflare 1.1.1.1 often <10 ms p50 (reported class)5L4 LB                 TCP / connection fan-out           HAProxy-class: 100k+ conn/box (heuristic/bench)6L7 LB / proxy         HTTP path/header route             Envoy/NGINX: tens of kRPS/core class (heuristic)7CDN                   edge cache static+dynamic          Cloudflare-scale: 300+ cities, 100+ Tbps capacity (reported marketing)8App server            stateless HTTP/gRPC                remote gRPC same-region ~5–20 ms (heuristic)9Cache Redis/MC        in-memory K/V                     Memcached ~1M ops/s/box class; Redis 100k–1M ops/s depending size (heuristic/bench)10Queue / log           Kafka / SQS / Rabbit               Stripe data stack: ~700 TB/day publish, ~50 Kafka clusters (reported 2024)11                                                         LinkedIn Kafka: trillions msgs/day class (reported historical)12OLTP Postgres/MySQL   row store ACID                    single primary ~10–50k writes/s class before serious sharding (heuristic)13Wide-col Cassandra    partition-key scale                Discord: 177 Cassandra nodes → 72 Scylla for trillions msgs (reported 2023)14Object store S3/GCS   blobs                             S3 multi-million req/s internal class (reported talks); cold GET p99 tens–100ms15Search ES/OpenSearch  inverted index                    single shard ~5–20k docs/s index; query ms–tens ms (heuristic)16Stream Flink/KStreams event-time jobs                   million events/s/cluster class (heuristic)1718Use numbers to force forks: if origin sees 30k RPS of 50 KB pages, you need a CDN — not a bigger Postgres.

Load balancing — L4 vs L7, health, draining

A load balancer spreads traffic across instances. L4 (TCP) is cheap and connection-aware (raw throughput, long-lived sockets, gRPC passthrough). L7 (HTTP) routes by path/header/cookie (canaries, sticky sessions, path split, WAF hooks). Know health checks and connection draining on deploy so in-flight requests finish. Consistent hashing / affinity keeps a key on one backend (session cache, WebSocket); plan for reshuffles when nodes join/leave (virtual nodes reduce movement). An LB does not fix a single-writer DB, global lock, or hot partition — always name the remaining choke point after you place the LB.
code
1L4 vs L7 TRADEOFF2                 L4 (TCP/NLB)              L7 (HTTP/ALB/Envoy)3Latency         lower overhead             more CPU per request4Routing         IP/port only               path, header, cookie, gRPC method5TLS             often passthrough          terminate + re-encrypt common6Sticky          5-tuple hash               cookie / header affinity7Use when        extreme conn count,        canaries, auth at edge, path split8                raw throughput9Does NOT fix    hot partition, single primary write, N+1 query storms

Caching — aside, through, behind, invalidation

Cache-aside (lazy): app reads cache; miss → DB → fill. Simple, app-owned, classic for objects/sessions. Write-through: cache+DB update together (fresher writes, higher write latency). Write-behind: ack from cache, flush DB async (fast writes, crash risk). Pick by read:write ratio, staleness budget, and source of truth. Invalidation: TTL, delete-on-write, or versioned keys. Stampede: single-flight / soft TTL so hot key expiry does not melt the DB. Senior signal: say what is stale-ok vs must-be-fresh before placing the cache. Redis vs Memcached: Memcached is multi-thread pure object cache at huge ops/s; Redis adds structures (sorted sets, streams), Lua atomics, optional persistence — pick by need for semantics, not fashion.
python
1# Cache-aside with single-flight (sketch)2def get_user(user_id):3	v = cache.get(f"user:{user_id}")4	if v is not None:5		return v6	with singleflight.lock(user_id):          # one DB load per key under stampede7		v = cache.get(f"user:{user_id}")8		if v is not None:9			return v10		v = db.query("SELECT * FROM users WHERE id=%s", user_id)11		cache.set(f"user:{user_id}", v, ttl=60)12		return v1314# On write: update DB first, then DELETE the key (not write-through),15# so the next read reloads the new row. Accept short race windows.
code
1CACHE EVICTION TRADEOFF2Policy      Best for                          Cost / caveat3----------  --------------------------------  ------------------------4LRU         general hot-key object cache      two pointers; scan risk5LFU         long-lived popularity (news)      counters; aging needed6TTL only    sessions, rate windows            no size bound alone7TinyLFU     mixed web workloads               more complex89SYNC vs ASYNC TRANSPORT10                 gRPC / REST                 Kafka / SQS11Latency         1 RTT, lower                  + durable hop, higher p5012Coupling        caller waits                  producer decoupled13Failure         caller sees 5xx               200 + consumer lag/fail14Replay          none                         hours → forever (log)15Use when        inside request boundary       cross-team, spike absorb

Senior signals vs mid-level failures (primitives)

code
1TOPIC        MID-LEVEL FAILURE                    SENIOR SIGNAL2-----------  -----------------------------------  -----------------------------------------3Cache        “Throw Redis in front of Postgres”   Names failure mode (hot key, fan-out collapse)4                                                   + eviction matching access pattern5Queue        “We need Kafka”                      Task queue (SQS) vs event log (Kafka);6                                                   replay vs work-item semantics7DB           “Postgres for everything”            Postgres for transactional core; wide-col for8                                                   partition-key chat/feed; S3 for cold analytics9CDN          “Use CloudFront”                     Static vs dynamic edge vs regional cache by10                                                   hit ratio and request fan-out11Shard        “Hash user_id, done”                 Hot-key plan (celebrity, viral code, whale tenant)

Sharding, replication, consistency

Hash shard by user/entity for even load; range for time scans (hot newest tail); geo/tenant for residency/noisy neighbors; consistent hashing to reduce remaps. Always name partition key, hot-key risk, rebalance. Celebrities, viral codes, whale tenants break “even” hashes — salt, celebrity tier, or pull path. Leader–follower replication: simple, lag is the tax. Multi-leader needs conflict rules. Own the vocabulary: strong, eventual, read-your-writes, causal — picked from a user-visible promise (“I see my edit now; others may lag 2s”), not from “CAP says.” Most apps never need sharded Postgres; prove write QPS exceeds one primary (~10–50k/s heuristic) before you inherit reshard and cross-shard join pain.
Hot partitions kill designs that look fine on paper. Senior answers do not stop at “we shard by user_id” — they name the power-law key and the mitigation (salting, separate tier, local limits, or changing fan-out for that entity). Discord’s message model (as reported) partitions by (channel_id, time_bucket) so one chatty channel cannot create an unbounded forever-hot row.

Queues, IDs, and idempotency

Queues absorb spikes, retry, batch, and protect interactive p99. Log/stream (Kafka-style): ordered partitions, replay, many consumer groups — Stripe’s data stack (reported) uses Kafka as a durable ordered replayable backbone at huge daily TB volume. Work queue (SQS-style): simple workers, less replay story. Default: at-least-once + idempotent consumers — “exactly once” is almost always effect-once via dedupe keys. Ops: DLQ, lag SLOs, backpressure, partition key for ordering. IDs: autoincrement (hard to shard writes), UUID (easy, non-sortable), ULID/Snowflake (time-sortable distributed), base62 (shortener). Stripe-style Idempotency-Key on writes stops double-charge on retry. “Throw a queue in front” is often wrong: you add at-least-once tax, higher e2e latency, worse observability, poison messages — ask whether you need durability+replay or just backpressure.
python
1# Idempotent worker pattern (at-least-once delivery)2def handle_event(event):3	# event.id is a stable business key from the producer4	if already_processed(event.id):5		return  # safe redelivery6	apply_side_effect(event)         # DB write, email, index update7	mark_processed(event.id)         # same transaction when possible89# Ordering: partition by channel_id so messages in one channel10# stay ordered; global order across all channels is usually not required.

Rate limit, CDN, search, realtime, ops (altitude)

Rate limit: fixed window (simple, boundary burst), sliding log (accurate, heavy), token bucket (burst + sustained — default; Stripe’s public rate-limit writing is a classic reported pattern with Redis + atomic scripts). Place at edge (coarse), gateway (per user/key), service (expensive ops). CDN: three buys — latency (bytes closer), origin offload (80–95% hit class when cacheable), DDoS absorption. It does not speed personalized DB reads if hit ratio is low. Search: inverted index is a derived view (async/CDC), not system of record. WebSockets: gateway maps user→conn, heartbeats, pub-sub fan-out; Slack’s public write-up (reported) on moving millions of concurrent WS toward Envoy is a good sticky-connection scale story. Observability/cost: RED/USE + traces; avoid unbounded metric labels; always leave a $/1k or storage-growth sentence. Multi-region depth is L3.
python
1# Inputs (heuristic interview scenario)2# 200M DAU; 5% hit trending; 50 KB page; treat as ~50 reads/active-user/day on that surface3dau = 200_000_0004active = dau * 0.055reads_day = active * 50              # 500M reads/day6avg_rps = reads_day / 86_400         # ~5_800 RPS7peak_rps = avg_rps * 5               # ~30_000 RPS peak (5× diurnal heuristic)8egress_peak = peak_rps * 50_000      # ~1.5 GB/s if origin served all bytes910# Decisions forced by the numbers:11# 1) CDN/full-page edge cache — 95% hit → origin ~1.5k RPS not 30k12# 2) Trending list precomputed (Flink/Kafka Streams) → Redis sorted set, not live SQL join13# 3) Redis failure mode: origin must shed load; size for partial fallback + backpressure14# Without 30k peak you might skip CDN; without 1.5 GB/s you might skip edge bodies.

Worked: annotate a request lifecycle

Walk this aloud until automatic. Example: user opens a product page with images and a reviews widget. Label each hop: cache pattern, consistency, failure mode, metric.
code
1BROWSER2  │  DNS / geo-DNS3  ▼4CDN (images, JS) ── hit? return ── miss? origin pull5  │6  ▼7EDGE / L7 LB ── TLS, WAF, coarse rate limit8  │9  ▼10API GATEWAY ── auth, per-user rate limit, routing11  │12  ▼13PRODUCT SERVICE14  ├─ cache-aside (product:{id}) ── hit? return15  │                              miss? PRIMARY DB (or replica if stale-ok)16  ├─ async: emit "product.viewed" → QUEUE → analytics / recsys workers17  └─ reviews: secondary call or BFF aggregation18  │19  ▼20SEARCH / FEED indexes  (async, not on the critical path for this read)2122Label each hop: cache pattern, consistency, failure mode, metric.
If you can narrate every hop — and name what fails when that hop is down — you already sound senior before the deep dive starts.

Interview answers — primitives (speak these cold)

  1. 01Redis vs Memcached? → Memcached multi-thread pure cache; Redis structures/Lua/persist — pick by semantics needed.
  2. 02Cassandra vs Postgres for chat? → Partition-key time-range reads + write scale → wide-col; joins/tx → Postgres.
  3. 03Sync vs async? → Sync inside request boundary; async across team/outage domains with idempotent consumers.
  4. 04Why not always a queue? → Latency, observability, at-least-once tax; need durability/replay or just backpressure?
  5. 05SQL vs NoSQL? → Start SQL until it hurts; pain type picks the store (write throughput, partition access, search).
  6. 06What does CDN buy? → Latency, origin offload, DDoS absorption — not free personalization.
  7. 07Shard Postgres already? → Prove write QPS needs it; replicas first; reshard/joins/hot shards are the tax.
  8. 08Ordering guarantee? → FIFO when money/correctness; best-effort analytics; blast radius of reorder.
  9. 09Where put the cache? → After stating staleness; cache-aside for app objects; never system of record.
  10. 10How do you shard? → Partition key, access pattern, hot-key plan, rebalance story.
  11. 11Exactly-once? → At-least-once + idempotent handlers + dedupe keys — not broker magic.
  12. 12Strong or eventual? → Restate user promise; read-your-writes for self; eventual for derived views.
System Design Primer — How to start with system designGaurav SenarticleDiscord — How Discord Stores Trillions of MessagesDiscord EngineeringarticleStripe — Idempotency (building-block write primitive)StripearticleCloudflare — What is a CDN?CloudflaredocsLatency Numbers Every Programmer Should Know (gist)Jeff Dean / jboner gistdocsSystem Design Primer (building blocks)donnemartin

Checkpoint

You have a product detail page: 100:1 read:write, can tolerate ~30s stale prices in a pinch, must never show another user’s cart. Best cache approach?

AWrite-behind cache as system of record for product + cartBCache-aside for public product HTML/JSON with TTL + invalidate on edit; never cache another user’s cart in a shared keyCNo cache — only strong reads from primary for everything
Sign up free to answer and see why

Checkpoint

Interview: “Users must immediately see their own profile edit, but other users can lag a few seconds.” Which consistency story is right?

AFull serializable isolation on every read worldwideBRead-your-writes for the editing user (sticky primary / version token); eventual for others via replica lagCEventual for everyone including the editor
Sign up free to answer and see why

Checkpoint

A shard key of user_id looks even, but p99 spikes whenever a celebrity posts. What is the real problem?

AThe load balancer is misconfiguredBHot partition / hot key — one user_id concentrates fan-out or write traffic; need salting, celebrity tier, or pull pathCNeed more replicas of every shard equally
Sign up free to answer and see why

Checkpoint

A read-heavy social app needs p99 reads < 5 ms for a 10 KB user profile, users are global, DB is Postgres in us-east-1. Best direction?

ASingle bigger Postgres instance in us-east-1 onlyBRegional caches (and/or regional read path) so the 10 KB hot object is served in-region; replicas help but pure remote primary cannot hit global 5 msCShard Postgres + Memcached in one region only
Sign up free to answer and see why

Checkpoint

In the annotated lifecycle, where should full-text search indexing usually run for a product create?

ASynchronously in the create request before 200 OKBAsynchronously via queue/CDC after the system-of-record write commitsCOnly at the CDN layer
Sign up free to answer and see why

Could you walk a request hop-by-hop and pick cache, shard, queue, and consistency patterns with a one-line tradeoff each?

New to itGetting thereConfident

Takeaways

  • ~12 primitives cover most classic SD rounds — pick by pressure, not by fashion.
  • Cache = pattern + staleness budget; invalidation and stampede matter more than the brand.
  • Shard key + hot-key plan is half the design; replication needs a user-visible consistency story.
  • Queues buy latency and retries; consumers must be idempotent; task queue ≠ event log.
  • Label production numbers reported vs heuristic; every estimate should force a design fork.

Next: the 45–60 minute framework — how seniors structure the round and turn capacity numbers into decisions.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.