Lesson 2 of 8 · 55 min
The 45–60 min framework + capacity with a why
Five-step interview OS: clarify, estimate with decision-linked math, API/model, deep dive, wrap — plus L4/L5/L6 bars and anti cargo-cult estimation.
Depth of reasoning, not diagram bingo
The 5-step framework (and why each step exists)
1STEP WHAT YOU DO WHY IT EXISTS2------------------ --------------------------------------- --------------------------------------31 Clarify (~5m) F/NFR, personas, scale vectors Half of mid failures = wrong problem42 Estimate (~5m) QPS peak, storage 5y, bandwidth Forces defensible tradeoffs53 Skeleton (~10m) Boxes: client→LB→API→DB/cache/queue Credibility; surfaces unknowns64 Deep dives (~25m) 2–3 high-risk components only Senior judgment = where risk lives75 Wrap-up (~10m) Failures, monitors, rejected options Explains reasoning, not just diagram89Steps 1 and 5 expose reasoning. Step 3 is coherence. Steps 2 and 4 are senior judgment.- 01Open: “I’ll confirm scope, estimate, then design the hot path and deep-dive the bottleneck.”
- 02Ask 4–6 questions max; batch them; don’t interrogate for 15 minutes.
- 03State non-goals: “No ML ranking in v1; chronological feed only.”
- 04Write NFRs as numbers or ranges, not vibes: “p99 read 200ms, durable writes, eventual feed OK.”
/ 86400, then apply a peak factor (often 2–5× for diurnal; events can be 10×). Storage: entities × size × retention × replication. Bandwidth: QPS × payload.Constant ladder (memorize small, derive the rest)
1# Rough ladder — label as INTERVIEW HEURISTICS, not measured prod metrics2SECONDS_PER_DAY = 86_400 # ~1e534# Traffic intuition anchors5# 1k QPS = small / early product6# 10k QPS = serious engineering starts7# 100k QPS = distributed systems required on hot path8# 1M QPS = hyper-scale class problem910# Payload intuition11TEXT_MSG = 1_000 # ~1 KB12PHOTO = 100_000 # ~100 KB13VIDEO_CLIP = 5_000_000 # ~5 MB1415# Example: 50M DAU, 20 feed reads/user/day, 2 writes/user/day16dau = 50_000_00017read_qps = dau * 20 / 86_400 # ~12K QPS avg18write_qps = dau * 2 / 86_400 # ~1.2K QPS avg19peak_read = read_qps * 5 # ~60K peak with 5× heuristic20# → cache/materialize reads; write path can stay simpler + async fan-outKey idea
Common mistake
“I need precise storage to three decimal places.”
Three estimation examples that force architecture
1EXAMPLE A — Timeline read (Twitter-class numbers are archival/heuristic)2 300M users, 10% DAU on home = 30M; 5 reads/user/day → ~150M reads/day3 → ~1,750 RPS avg → ~10k RPS peak (×5–6)4 DECISION: at 10k RPS peak you cannot afford pure fan-out-on-read for everyone5 → push/precompute timelines for normal users (see L5)67EXAMPLE B — URL shortener redirects8 200M clicks/day → ~2,300 RPS avg → ~10k peak9 DECISION: hash lookup must be memory/edge; cold disk too slow for p99 redirect1011EXAMPLE C — Chat messages (Discord-class shape, illustrative)12 1B users × 10% senders × 50 msgs/day → 5B msgs/day → ~58k msg/s avg → peak much higher13 DECISION: partitioned append-only store by (channel, time_bucket), not unpartitioned Postgres1415READ-HEAVY vs WRITE-HEAVY (quick tradeoff)16 Read-heavy (timeline) Write-heavy (chat)17Cache edge + materialized views skip heavy cache; write path first18DB Postgres + replicas often OK wide-col / log-oriented19Hot spot celebrity cache miss active channel write hotspot20Cost driver egress + fan-out storage storage IOPS + retention1# URL shortener sketch (preview L4) — model first2# links(short_code PK, long_url, owner_id, created_at, expires_at)3# analytics_events(short_code, ts, country, ua) -- append-only, async45# POST /v1/links {long_url, custom_alias?} -> {short_url}6# GET /{code} -> 302 Location: long_url (hot path)7# Idempotency-Key on POST to avoid double creates on retryL4 vs L5 vs L6 expectations
Senior signal is the inverse of the #1 reject reason: do not jump to boxes before requirements, estimates, and data model — and do not finish without failure modes and a cost or scale ceiling sentence.
Anti cargo-cult estimation
- 01Derive from the prompt’s DAU/usage — don’t paste Twitter numbers into a ride-sharing prompt.
- 02Show units: users/day × actions/user ÷ 86400 = QPS.
- 03Peak factor + growth (“what if 10×?”) beats false precision.
- 04Tie each number to a fork: shard count, cache, retention, replica count.
- 05Refuse: “and then Kafka + Redis + Cassandra” with no why.
- 06State assumptions out loud so the interviewer can correct them.
- 07Round aggressively: 1 sig fig is a feature.
- 08Stop estimating when the design constrains itself (“100 PB → object store”).
Common mistake
“If I memorize enough production case studies I’ll pass.”
When you get stuck — order of operations
1# Decision checklist you can run mentally in 30s2# 1. Read-heavy or write-heavy?3# 2. Latency budget on the interactive path?4# 3. Consistency: self / global / derived?5# 4. Fan-out: who explodes — writers or readers?6# 5. Stateful connection or request/response?7# 6. What is async-ok?8# 7. What is the hot key?9# 8. What metric pages the on-call?10# 9. What would I NOT build at this scale?11# 10. What changes at 10×?Key idea
Interview answers — framework & capacity
- 01Estimate QPS? → DAU × actions/day ÷ 86400 → peak ×5–10 → read/write split → growth ×3.
- 02Numbers to memorize? → latency ladder, ~1M cache ops/s/box class, ~10k RDS writes/s class, Kafka per-partition throughput class.
- 03How to estimate without lying? → state assumptions, anchors 1k/10k/100k/1M/1B, show arithmetic.
- 04Why peak=5× avg? → diurnal; sports/elections/news can be 10×.
- 05Cost of 10× wrong? → traffic 10× but cache working set can blow 100×.
- 065y storage? → daily bytes × 365 × 5 × replication (often 3×) × backup factor.
- 07When stop estimating? → when number forces architecture (PB → object store).
- 08Most common mistake? → unit errors (MB/GB, s/ms) and forgetting peak vs average.
- 09L5 vs L4? → drives scope, alternatives+default, failure/cost unprompted.
- 10Stuck with 40 min left? → freeze MVP, sketch hot path, deep-dive bottleneck.
System Design Interview — Step By Step GuideByteByteGoarticleinterviewing.io — System Design Interview Guide for Senior Engineersinterviewing.ioarticleBack-of-the-envelope estimation (ByteByteGo)ByteByteGodocsJeff Dean latency numbers (gist)jboner / DeanarticleHello Interview — System Design in a HurryHello InterviewarticleByteByteGo — Framework for system design interviewsByteByteGoWorked capacity walk-through you can say aloud
1SPOKEN TEMPLATE (60s capacity)21. DAU and core action rate32. Avg QPS = DAU × actions / 8640043. Peak = avg × 3–10 (state which)54. Read:write split65. Payload × QPS = bandwidth76. Storage = entities × size × years × RF87. THEREFORE: [cache | shard | async | CDN | hybrid fanout]910If step 7 is missing, the math was theater.How to deep-dive without drowning
1DEEP-DIVE SCORECARD (self-check)2[ ] Named the riskiest component with a reason3[ ] Showed schema or API fields for that component4[ ] Gave 2 alternatives + default5[ ] Named a failure mode + mitigation6[ ] Named a metric / SLO for that component7[ ] Said what you would not build yet8[ ] Left 5 min for wrap-upCheckpoint
You have 40 minutes left and still no diagram. You feel stuck on edge cases. Best move?
Checkpoint
Your estimate shows ~15K read QPS and ~200 write QPS for a feed. Which design decision does that number most directly force?
Checkpoint
What best distinguishes an L5 answer from an L4 answer in the same prompt?
Checkpoint
Candidate computes storage as 847.32 TB exactly from stacked assumptions. Interviewer looks unimpressed. Why?
Checkpoint
You estimate 50k RPS peak and provision 3 instances. At 3 AM p99 is 700 ms. What did the capacity plan most likely forget?
Worked math card — keep until automatic
1CAPACITY CARD (fill blanks in every mock)2DAU = ________ actions/user/day = ________3avg QPS = DAU * actions / 86400 = ________4peak QPS = avg * peak_factor(3..10) = ________5read QPS = ________ write QPS = ________6payload bytes = ________ bandwidth = peak*payload = ________7storage/year = entities/day * size * 365 * RF(3) = ________8THEREFORE (must name a box change):9 [ ] cache [ ] shard [ ] async queue10 [ ] CDN [ ] hybrid fanout [ ] cold tier11 [ ] multi-region read [ ] rate limit expensive ops1213LEVEL BAR REMINDER14L4: right boxes with help; basic estimate15L5: drives clock; alternatives+default; failures unprompted16L6: + multi-region, unit cost, cardinality, rollout/on-call1718ANTI-PATTERNS19- Twitter numbers pasted into ride-hail prompt20- 847.32 TB precision without a decision21- 12 microservices before API/schema22- Silent diagram polishing for 20 minutes23- Skipping failure modes in the last 5 minutes2425CLOCK (60m)260-8 clarify F/NFR + non-goals278-13 capacity with THEREFORE2813-20 API + data model2920-35 high-level + pick deep dive3035-50 deep dive riskiest component3150-60 failures, metrics, cost, rejected optionsCan you run a 45–60 min SD round structure cold, with capacity math that forces at least one design fork?
Takeaways
- Rubric = framing, decision-linked capacity, data model, deep dive, failure/cost — not logo density.
- Time-box: clarify → estimate → API/model → high-level → deep dive → wrap.
- L5 drives; L6 adds multi-region, cost, ops, cardinality.
- Constant ladder + peak factor + one sig fig; every number changes a box.
- When stuck: shrink scope, pick the bottleneck, narrate tradeoffs.
Next: the senior layer — multi-region, multi-tenant isolation, cost, and ops that separate L5 from L6.
Sources
Free to read · better with Enzo
Learn it with Enzo
Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.