Lesson 1 of 8 · 55 min
Building blocks every round reuses
Load balancers, caches, sharding, queues, IDs, CDN, search, WebSockets, observability — with reported/heuristic production numbers and a full request lifecycle.
Stop inventing boxes from scratch
Latency ladder (memorize once, reuse forever)
1LATENCY LADDER (heuristic — interview use, not a SLA)2 L1 cache hit ~1 ns3 Branch mispredict ~3 ns4 L2 hit ~5 ns5 Mutex lock/unlock ~17 ns6 Main memory ~100 ns7 SSD random read ~100 us (1e5 ns)8 Same datacenter RTT ~0.5 ms9 SSD / NVMe sequential higher throughput, still not free10 Disk seek ~1–10 ms11 Same region cross-AZ ~1–2 ms class12 Cross-city ~5–20 ms13 Cross-ocean ~50–150 ms1415DESIGN RULE: p99 budget ÷ RTT = how many remote hops you can afford.16If p99=100ms and hop=5ms, you do not get 30 sequential service calls.Production building-block numbers (reported / heuristic)
1BLOCK WHAT IT IS NUMBER (label)2-------------------- --------------------------------- ----------------------------------3DNS hierarchical lookup cold multi-continent ~50–120 ms (heuristic);4 Cloudflare 1.1.1.1 often <10 ms p50 (reported class)5L4 LB TCP / connection fan-out HAProxy-class: 100k+ conn/box (heuristic/bench)6L7 LB / proxy HTTP path/header route Envoy/NGINX: tens of kRPS/core class (heuristic)7CDN edge cache static+dynamic Cloudflare-scale: 300+ cities, 100+ Tbps capacity (reported marketing)8App server stateless HTTP/gRPC remote gRPC same-region ~5–20 ms (heuristic)9Cache Redis/MC in-memory K/V Memcached ~1M ops/s/box class; Redis 100k–1M ops/s depending size (heuristic/bench)10Queue / log Kafka / SQS / Rabbit Stripe data stack: ~700 TB/day publish, ~50 Kafka clusters (reported 2024)11 LinkedIn Kafka: trillions msgs/day class (reported historical)12OLTP Postgres/MySQL row store ACID single primary ~10–50k writes/s class before serious sharding (heuristic)13Wide-col Cassandra partition-key scale Discord: 177 Cassandra nodes → 72 Scylla for trillions msgs (reported 2023)14Object store S3/GCS blobs S3 multi-million req/s internal class (reported talks); cold GET p99 tens–100ms15Search ES/OpenSearch inverted index single shard ~5–20k docs/s index; query ms–tens ms (heuristic)16Stream Flink/KStreams event-time jobs million events/s/cluster class (heuristic)1718Use numbers to force forks: if origin sees 30k RPS of 50 KB pages, you need a CDN — not a bigger Postgres.Load balancing — L4 vs L7, health, draining
1L4 vs L7 TRADEOFF2 L4 (TCP/NLB) L7 (HTTP/ALB/Envoy)3Latency lower overhead more CPU per request4Routing IP/port only path, header, cookie, gRPC method5TLS often passthrough terminate + re-encrypt common6Sticky 5-tuple hash cookie / header affinity7Use when extreme conn count, canaries, auth at edge, path split8 raw throughput9Does NOT fix hot partition, single primary write, N+1 query stormsCommon mistake
“Just put an ALB in front and we’re scalable.”
Caching — aside, through, behind, invalidation
1# Cache-aside with single-flight (sketch)2def get_user(user_id):3 v = cache.get(f"user:{user_id}")4 if v is not None:5 return v6 with singleflight.lock(user_id): # one DB load per key under stampede7 v = cache.get(f"user:{user_id}")8 if v is not None:9 return v10 v = db.query("SELECT * FROM users WHERE id=%s", user_id)11 cache.set(f"user:{user_id}", v, ttl=60)12 return v1314# On write: update DB first, then DELETE the key (not write-through),15# so the next read reloads the new row. Accept short race windows.1CACHE EVICTION TRADEOFF2Policy Best for Cost / caveat3---------- -------------------------------- ------------------------4LRU general hot-key object cache two pointers; scan risk5LFU long-lived popularity (news) counters; aging needed6TTL only sessions, rate windows no size bound alone7TinyLFU mixed web workloads more complex89SYNC vs ASYNC TRANSPORT10 gRPC / REST Kafka / SQS11Latency 1 RTT, lower + durable hop, higher p5012Coupling caller waits producer decoupled13Failure caller sees 5xx 200 + consumer lag/fail14Replay none hours → forever (log)15Use when inside request boundary cross-team, spike absorbKey idea
Senior signals vs mid-level failures (primitives)
1TOPIC MID-LEVEL FAILURE SENIOR SIGNAL2----------- ----------------------------------- -----------------------------------------3Cache “Throw Redis in front of Postgres” Names failure mode (hot key, fan-out collapse)4 + eviction matching access pattern5Queue “We need Kafka” Task queue (SQS) vs event log (Kafka);6 replay vs work-item semantics7DB “Postgres for everything” Postgres for transactional core; wide-col for8 partition-key chat/feed; S3 for cold analytics9CDN “Use CloudFront” Static vs dynamic edge vs regional cache by10 hit ratio and request fan-out11Shard “Hash user_id, done” Hot-key plan (celebrity, viral code, whale tenant)Sharding, replication, consistency
(channel_id, time_bucket) so one chatty channel cannot create an unbounded forever-hot row.Common mistake
“Eventual consistency means the system is broken sometimes.”
Queues, IDs, and idempotency
1# Idempotent worker pattern (at-least-once delivery)2def handle_event(event):3 # event.id is a stable business key from the producer4 if already_processed(event.id):5 return # safe redelivery6 apply_side_effect(event) # DB write, email, index update7 mark_processed(event.id) # same transaction when possible89# Ordering: partition by channel_id so messages in one channel10# stay ordered; global order across all channels is usually not required.Key idea
Rate limit, CDN, search, realtime, ops (altitude)
Capacity example that forces design (trending page)
1# Inputs (heuristic interview scenario)2# 200M DAU; 5% hit trending; 50 KB page; treat as ~50 reads/active-user/day on that surface3dau = 200_000_0004active = dau * 0.055reads_day = active * 50 # 500M reads/day6avg_rps = reads_day / 86_400 # ~5_800 RPS7peak_rps = avg_rps * 5 # ~30_000 RPS peak (5× diurnal heuristic)8egress_peak = peak_rps * 50_000 # ~1.5 GB/s if origin served all bytes910# Decisions forced by the numbers:11# 1) CDN/full-page edge cache — 95% hit → origin ~1.5k RPS not 30k12# 2) Trending list precomputed (Flink/Kafka Streams) → Redis sorted set, not live SQL join13# 3) Redis failure mode: origin must shed load; size for partial fallback + backpressure14# Without 30k peak you might skip CDN; without 1.5 GB/s you might skip edge bodies.Worked: annotate a request lifecycle
1BROWSER2 │ DNS / geo-DNS3 ▼4CDN (images, JS) ── hit? return ── miss? origin pull5 │6 ▼7EDGE / L7 LB ── TLS, WAF, coarse rate limit8 │9 ▼10API GATEWAY ── auth, per-user rate limit, routing11 │12 ▼13PRODUCT SERVICE14 ├─ cache-aside (product:{id}) ── hit? return15 │ miss? PRIMARY DB (or replica if stale-ok)16 ├─ async: emit "product.viewed" → QUEUE → analytics / recsys workers17 └─ reviews: secondary call or BFF aggregation18 │19 ▼20SEARCH / FEED indexes (async, not on the critical path for this read)2122Label each hop: cache pattern, consistency, failure mode, metric.If you can narrate every hop — and name what fails when that hop is down — you already sound senior before the deep dive starts.
Interview answers — primitives (speak these cold)
- 01Redis vs Memcached? → Memcached multi-thread pure cache; Redis structures/Lua/persist — pick by semantics needed.
- 02Cassandra vs Postgres for chat? → Partition-key time-range reads + write scale → wide-col; joins/tx → Postgres.
- 03Sync vs async? → Sync inside request boundary; async across team/outage domains with idempotent consumers.
- 04Why not always a queue? → Latency, observability, at-least-once tax; need durability/replay or just backpressure?
- 05SQL vs NoSQL? → Start SQL until it hurts; pain type picks the store (write throughput, partition access, search).
- 06What does CDN buy? → Latency, origin offload, DDoS absorption — not free personalization.
- 07Shard Postgres already? → Prove write QPS needs it; replicas first; reshard/joins/hot shards are the tax.
- 08Ordering guarantee? → FIFO when money/correctness; best-effort analytics; blast radius of reorder.
- 09Where put the cache? → After stating staleness; cache-aside for app objects; never system of record.
- 10How do you shard? → Partition key, access pattern, hot-key plan, rebalance story.
- 11Exactly-once? → At-least-once + idempotent handlers + dedupe keys — not broker magic.
- 12Strong or eventual? → Restate user promise; read-your-writes for self; eventual for derived views.
System Design Primer — How to start with system designGaurav SenarticleDiscord — How Discord Stores Trillions of MessagesDiscord EngineeringarticleStripe — Idempotency (building-block write primitive)StripearticleCloudflare — What is a CDN?CloudflaredocsLatency Numbers Every Programmer Should Know (gist)Jeff Dean / jboner gistdocsSystem Design Primer (building blocks)donnemartinCheckpoint
You have a product detail page: 100:1 read:write, can tolerate ~30s stale prices in a pinch, must never show another user’s cart. Best cache approach?
Checkpoint
Interview: “Users must immediately see their own profile edit, but other users can lag a few seconds.” Which consistency story is right?
Checkpoint
A shard key of user_id looks even, but p99 spikes whenever a celebrity posts. What is the real problem?
Checkpoint
A read-heavy social app needs p99 reads < 5 ms for a 10 KB user profile, users are global, DB is Postgres in us-east-1. Best direction?
Checkpoint
In the annotated lifecycle, where should full-text search indexing usually run for a product create?
Could you walk a request hop-by-hop and pick cache, shard, queue, and consistency patterns with a one-line tradeoff each?
Takeaways
- ~12 primitives cover most classic SD rounds — pick by pressure, not by fashion.
- Cache = pattern + staleness budget; invalidation and stampede matter more than the brand.
- Shard key + hot-key plan is half the design; replication needs a user-visible consistency story.
- Queues buy latency and retries; consumers must be idempotent; task queue ≠ event log.
- Label production numbers reported vs heuristic; every estimate should force a design fork.
Next: the 45–60 minute framework — how seniors structure the round and turn capacity numbers into decisions.
Sources
Free to read · better with Enzo
Learn it with Enzo
Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.