Lesson 2 of 8 · 55 min

The 45–60 min framework + capacity with a why

Five-step interview OS: clarify, estimate with decision-linked math, API/model, deep dive, wrap — plus L4/L5/L6 bars and anti cargo-cult estimation.

Depth of reasoning, not diagram bingo

Interviewers grade process under uncertainty: scope, estimates tied to choices, right primitives, tradeoffs spoken aloud, failure modes, cost. A pretty boxes-and-arrows sketch that never names what would change at 10× is a mid-level performance. This lesson is the operating system for every worked design that follows. The interviewer is not grading whether you can draw a system — they grade how you think when half the requirements are unknown. Rehearse the rejected-alternatives sentence every time: “I considered pure pull for all users but rejected it because peak home RPS makes fan-in too expensive; hybrid costs more engineering but caps write amplification.” Interviewers remember that sentence more than your boxes.
Default time budget for a ~60 minute round (adapt to 45 by shrinking deep-dive): 5–10 min requirements + NFRs; ~5 min capacity; 10–15 min high-level design; 20–25 min deep dive on the hard sub-problem; ~5–10 min failure modes, ops, cost, and rejected alternatives. If you burn 25 minutes drawing every microservice before clarifying writes vs reads, you fail on process even if the boxes are “correct.”

The 5-step framework (and why each step exists)

code
1STEP                 WHAT YOU DO                              WHY IT EXISTS2------------------  ---------------------------------------  --------------------------------------31 Clarify (~5m)     F/NFR, personas, scale vectors           Half of mid failures = wrong problem42 Estimate (~5m)    QPS peak, storage 5y, bandwidth          Forces defensible tradeoffs53 Skeleton (~10m)   Boxes: client→LB→API→DB/cache/queue      Credibility; surfaces unknowns64 Deep dives (~25m) 2–3 high-risk components only            Senior judgment = where risk lives75 Wrap-up (~10m)    Failures, monitors, rejected options     Explains reasoning, not just diagram89Steps 1 and 5 expose reasoning. Step 3 is coherence. Steps 2 and 4 are senior judgment.
Phase 1 — Requirements you drive. Restate the product in one sentence, then ask focused questions. Functional: core user actions, MVP vs out of scope, read vs write paths, online vs offline. Non-functional: QPS scale, latency SLO (p50/p99), durability, consistency, availability target, multi-region?, compliance/data residency?, cost sensitivity. Write assumptions on the board: “Assuming 50M DAU, 100:1 read:write, p99 < 200ms for reads, US-only MVP.” Senior signal: push back. “Do we need E2E encryption in MVP?” “Is global search required day one?” “Is this B2C viral or B2B 200 tenants?” Mid candidates wait to be guided; seniors scope and name non-goals. Staff+ candidates often reframe: “The hard problem here is fan-out, not auth — I’ll de-risk that first.” Ask 4–6 questions max; batch them; do not interrogate for 15 minutes.
  1. 01Open: “I’ll confirm scope, estimate, then design the hot path and deep-dive the bottleneck.”
  2. 02Ask 4–6 questions max; batch them; don’t interrogate for 15 minutes.
  3. 03State non-goals: “No ML ranking in v1; chronological feed only.”
  4. 04Write NFRs as numbers or ranges, not vibes: “p99 read 200ms, durable writes, eventual feed OK.”
Phase 2 — Capacity with a why (not theater). Capacity estimation exists to force design decisions, not to impress with arithmetic. Show order-of-magnitude math, 1 significant figure, then say “therefore we need X.” Cargo-cult failure: reciting “Twitter is 200M DAU and 100K QPS” without using a single number in a choice. Always convert DAU → average QPS with / 86400, then apply a peak factor (often 2–5× for diurnal; events can be 10×). Storage: entities × size × retention × replication. Bandwidth: QPS × payload.

Constant ladder (memorize small, derive the rest)

python Capacity anchors to keep on your mental whiteboard: 1k QPS is still mostly vertical scale; 10k QPS is where caches and async become default answers; 100k QPS is multi-service distributed design on the hot path; 1M QPS is specialized data planes. Storage anchors: 1 TB/year is one beefy primary-class problem; 100 TB forces tiering and sharding stories; multi-PB forces object store + metadata plane. Bandwidth anchors: 1 Gbps egress is a serious bill; 10 Gbps almost always means CDN or regionalization. Speak these ladders so estimates take seconds, not minutes.
1# Rough ladder — label as INTERVIEW HEURISTICS, not measured prod metrics2SECONDS_PER_DAY = 86_400  # ~1e534# Traffic intuition anchors5# 1k QPS   = small / early product6# 10k QPS  = serious engineering starts7# 100k QPS = distributed systems required on hot path8# 1M QPS   = hyper-scale class problem910# Payload intuition11TEXT_MSG   = 1_000       # ~1 KB12PHOTO      = 100_000     # ~100 KB13VIDEO_CLIP = 5_000_000   # ~5 MB1415# Example: 50M DAU, 20 feed reads/user/day, 2 writes/user/day16dau = 50_000_00017read_qps  = dau * 20 / 86_400   # ~12K QPS avg18write_qps = dau * 2  / 86_400   # ~1.2K QPS avg19peak_read = read_qps * 5        # ~60K peak with 5× heuristic20# → cache/materialize reads; write path can stay simpler + async fan-out

Three estimation examples that force architecture

code
1EXAMPLE A — Timeline read (Twitter-class numbers are archival/heuristic)2  300M users, 10% DAU on home = 30M; 5 reads/user/day → ~150M reads/day3  → ~1,750 RPS avg → ~10k RPS peak (×5–6)4  DECISION: at 10k RPS peak you cannot afford pure fan-out-on-read for everyone5            → push/precompute timelines for normal users (see L5)67EXAMPLE B — URL shortener redirects8  200M clicks/day → ~2,300 RPS avg → ~10k peak9  DECISION: hash lookup must be memory/edge; cold disk too slow for p99 redirect1011EXAMPLE C — Chat messages (Discord-class shape, illustrative)12  1B users × 10% senders × 50 msgs/day → 5B msgs/day → ~58k msg/s avg → peak much higher13  DECISION: partitioned append-only store by (channel, time_bucket), not unpartitioned Postgres1415READ-HEAVY vs WRITE-HEAVY (quick tradeoff)16                 Read-heavy (timeline)           Write-heavy (chat)17Cache            edge + materialized views       skip heavy cache; write path first18DB               Postgres + replicas often OK    wide-col / log-oriented19Hot spot         celebrity cache miss            active channel write hotspot20Cost driver      egress + fan-out storage        storage IOPS + retention
Phase 3 — API & data model before pretty topology. Contract-first: core endpoints, request/response fields, idempotency keys on writes. Data model: tables/collections, primary keys, secondary indexes, partition keys. Many seniors lose points by drawing 12 services with no schema — graders cannot tell if you understand access patterns. Access patterns drive schema, not ER purity.
code
1# URL shortener sketch (preview L4) — model first2# links(short_code PK, long_url, owner_id, created_at, expires_at)3# analytics_events(short_code, ts, country, ua)  -- append-only, async45# POST /v1/links {long_url, custom_alias?} -> {short_url}6# GET  /{code} -> 302 Location: long_url   (hot path)7# Idempotency-Key on POST to avoid double creates on retry
Phase 4 — High-level design. Client → edge/CDN → LB → API → services → storage → async workers. Identify the hard sub-problem early: fan-out, hot key, geo match, ordering, search freshness. Do not deep-dive auth unless asked. Name 2–3 alternatives for the hard part and pick a default: “Hybrid fan-out because pure write-path explodes for celebrities.” Phase 5 — Deep dive + failure + cost. Spend the bulk of the remaining time where the system breaks: thundering herd, partial outage, duplicate messages, clock skew, partition rebalance, multi-device delivery. Close with monitoring (what pages?), rollback, and rough cost or scaling ceilings. Staff+ adds multi-region and org/ops framing. Explicitly name trade-offs you rejected and why — that sentence alone often separates L5 from L4.

L4 vs L5 vs L6 expectations

L4 / mid: right primitives with guidance; clean diagram; basic estimates. L5 / senior: drives the round; 2–3 alternatives with a chosen default; tradeoffs in latency/throughput/cost/ops; failure modes unprompted. L6+ / staff: all of that plus multi-region, cost/finops, schema evolution, observability cardinality, rollout/on-call, sometimes org boundaries. Calibrate depth to the level on the req — do not monologue CRDTs for an L4 shortener.
Senior signal is the inverse of the #1 reject reason: do not jump to boxes before requirements, estimates, and data model — and do not finish without failure modes and a cost or scale ceiling sentence.

Anti cargo-cult estimation

  1. 01Derive from the prompt’s DAU/usage — don’t paste Twitter numbers into a ride-sharing prompt.
  2. 02Show units: users/day × actions/user ÷ 86400 = QPS.
  3. 03Peak factor + growth (“what if 10×?”) beats false precision.
  4. 04Tie each number to a fork: shard count, cache, retention, replica count.
  5. 05Refuse: “and then Kafka + Redis + Cassandra” with no why.
  6. 06State assumptions out loud so the interviewer can correct them.
  7. 07Round aggressively: 1 sig fig is a feature.
  8. 08Stop estimating when the design constrains itself (“100 PB → object store”).

When you get stuck — order of operations

1) Re-read the user journey on the board. 2) Ask which NFR is non-negotiable. 3) Shrink scope (MVP path only). 4) Pick the single bottleneck and deep-dive it. 5) Narrate tradeoffs even if the design is incomplete — incomplete-but-reasoned beats silent polishing of boxes.
code
1# Decision checklist you can run mentally in 30s2# 1. Read-heavy or write-heavy?3# 2. Latency budget on the interactive path?4# 3. Consistency: self / global / derived?5# 4. Fan-out: who explodes — writers or readers?6# 5. Stateful connection or request/response?7# 6. What is async-ok?8# 7. What is the hot key?9# 8. What metric pages the on-call?10# 9. What would I NOT build at this scale?11# 10. What changes at 10×?

Interview answers — framework & capacity

  1. 01Estimate QPS? → DAU × actions/day ÷ 86400 → peak ×5–10 → read/write split → growth ×3.
  2. 02Numbers to memorize? → latency ladder, ~1M cache ops/s/box class, ~10k RDS writes/s class, Kafka per-partition throughput class.
  3. 03How to estimate without lying? → state assumptions, anchors 1k/10k/100k/1M/1B, show arithmetic.
  4. 04Why peak=5× avg? → diurnal; sports/elections/news can be 10×.
  5. 05Cost of 10× wrong? → traffic 10× but cache working set can blow 100×.
  6. 065y storage? → daily bytes × 365 × 5 × replication (often 3×) × backup factor.
  7. 07When stop estimating? → when number forces architecture (PB → object store).
  8. 08Most common mistake? → unit errors (MB/GB, s/ms) and forgetting peak vs average.
  9. 09L5 vs L4? → drives scope, alternatives+default, failure/cost unprompted.
  10. 10Stuck with 40 min left? → freeze MVP, sketch hot path, deep-dive bottleneck.
System Design Interview — Step By Step GuideByteByteGoarticleinterviewing.io — System Design Interview Guide for Senior Engineersinterviewing.ioarticleBack-of-the-envelope estimation (ByteByteGo)ByteByteGodocsJeff Dean latency numbers (gist)jboner / DeanarticleHello Interview — System Design in a HurryHello InterviewarticleByteByteGo — Framework for system design interviewsByteByteGo

Worked capacity walk-through you can say aloud

Pick a prompt: “Design Instagram home for 200M DAU.” Speak: “Assume 200M DAU, 40% open home daily, 8 opens/user/day → 640M reads/day ≈ 7.4k RPS avg, peak ×5 ≈ 37k RPS. Posts 0.3/user/day among DAU → ~700 RPS avg posts. At 200 avg followers pure push ≈ 140k timeline writes/s — too high without hybrid. Therefore: push for normal, pull for celebs, timeline IDs + hydrate cache, CDN for media. Storage of timelines: 200M × 500 IDs × 16B ≈ 1.6 TB raw metadata class before replication — cluster, not one box.” Every sentence was a number forcing a box.
code
1SPOKEN TEMPLATE (60s capacity)21. DAU and core action rate32. Avg QPS = DAU × actions / 8640043. Peak = avg × 3–10 (state which)54. Read:write split65. Payload × QPS = bandwidth76. Storage = entities × size × years × RF87. THEREFORE: [cache | shard | async | CDN | hybrid fanout]910If step 7 is missing, the math was theater.
Common interviewer pushback: “Your peak factor is made up.” Answer: “Yes — it is an explicit assumption. If peaks are only 2× I still need cache; if they are 20× I need more aggressive edge and load shed. Tell me your peak model and I’ll recompute.” That response is senior: assumptions are negotiable; decision-linking is not.

How to deep-dive without drowning

Deep dive means one or two components for 15–25 minutes: data model fields, failure modes, scaling of that component, and tradeoffs rejected. Bad deep dive: listing every AWS service in the diagram. Good deep dive: “The fan-out worker: job schema, idempotency key, chunk size, retry/DLQ, lag SLO, what happens when celebrity threshold is wrong.” Invite the interviewer: “I’d go deep on fan-out — does that match what you want?”
code
1DEEP-DIVE SCORECARD (self-check)2[ ] Named the riskiest component with a reason3[ ] Showed schema or API fields for that component4[ ] Gave 2 alternatives + default5[ ] Named a failure mode + mitigation6[ ] Named a metric / SLO for that component7[ ] Said what you would not build yet8[ ] Left 5 min for wrap-up
Staff+ flavor in the same framework: add multi-region write topology, unit cost, schema evolution, and on-call pages — but only after the MVP hot path is solid. Jumping to global CRDTs before basic capacity is a mid-level tell wearing staff vocabulary.

Checkpoint

You have 40 minutes left and still no diagram. You feel stuck on edge cases. Best move?

AKeep asking clarifying questions until the design is fully specifiedBFreeze MVP + NFRs, sketch hot path + data model, then deep-dive the main bottleneckCStart drawing every microservice you know from a blog post
Sign up free to answer and see why

Checkpoint

Your estimate shows ~15K read QPS and ~200 write QPS for a feed. Which design decision does that number most directly force?

AUse blockchain for auditabilityBAvoid primary-DB hits on every read — cache and/or precompute timelines; keep write path simplerCSwitch the write path to multi-region active-active immediately
Sign up free to answer and see why

Checkpoint

What best distinguishes an L5 answer from an L4 answer in the same prompt?

AL5 draws more boxes and names more AWS servicesBL5 drives scope, proposes alternatives with a default, and discusses failure/cost without being promptedCL5 refuses to estimate because estimates are fake
Sign up free to answer and see why

Checkpoint

Candidate computes storage as 847.32 TB exactly from stacked assumptions. Interviewer looks unimpressed. Why?

AThey should have used binary TB not decimalBFalse precision + no link from the number to a design choice (tiering, retention, shard count)CStorage estimates are never allowed in interviews
Sign up free to answer and see why

Checkpoint

You estimate 50k RPS peak and provision 3 instances. At 3 AM p99 is 700 ms. What did the capacity plan most likely forget?

AA queue always fixes p99BPeak factor + redundancy headroom — 3 instances cannot absorb peak and survive one failureCThat Kubernetes autoscaling removes the need to think
Sign up free to answer and see why

Worked math card — keep until automatic

code
1CAPACITY CARD (fill blanks in every mock)2DAU = ________   actions/user/day = ________3avg QPS = DAU * actions / 86400 = ________4peak QPS = avg * peak_factor(3..10) = ________5read QPS = ________   write QPS = ________6payload bytes = ________   bandwidth = peak*payload = ________7storage/year = entities/day * size * 365 * RF(3) = ________8THEREFORE (must name a box change):9  [ ] cache     [ ] shard     [ ] async queue10  [ ] CDN       [ ] hybrid fanout     [ ] cold tier11  [ ] multi-region read     [ ] rate limit expensive ops1213LEVEL BAR REMINDER14L4: right boxes with help; basic estimate15L5: drives clock; alternatives+default; failures unprompted16L6: + multi-region, unit cost, cardinality, rollout/on-call1718ANTI-PATTERNS19- Twitter numbers pasted into ride-hail prompt20- 847.32 TB precision without a decision21- 12 microservices before API/schema22- Silent diagram polishing for 20 minutes23- Skipping failure modes in the last 5 minutes2425CLOCK (60m)260-8 clarify F/NFR + non-goals278-13 capacity with THEREFORE2813-20 API + data model2920-35 high-level + pick deep dive3035-50 deep dive riskiest component3150-60 failures, metrics, cost, rejected options

Can you run a 45–60 min SD round structure cold, with capacity math that forces at least one design fork?

New to itGetting thereConfident

Takeaways

  • Rubric = framing, decision-linked capacity, data model, deep dive, failure/cost — not logo density.
  • Time-box: clarify → estimate → API/model → high-level → deep dive → wrap.
  • L5 drives; L6 adds multi-region, cost, ops, cardinality.
  • Constant ladder + peak factor + one sig fig; every number changes a box.
  • When stuck: shrink scope, pick the bottleneck, narrate tradeoffs.

Next: the senior layer — multi-region, multi-tenant isolation, cost, and ops that separate L5 from L6.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.