Lesson 4 of 8 · 55 min
Worked design — URL shortener
Full shortener: base62/IDs, write and redirect paths, cache, hot aliases, async analytics, and estimation that kills over-engineering or forces edge cache.
Lesson 4 · Worked design
URL shortener end to end
Entry prompt, senior execution
Requirements drill
- 01Custom aliases? uniqueness scope?
- 02Expiry / one-time links?
- 03Auth for create? public create?
- 04Analytics real-time or daily?
- 05Update destination after create?
- 06Abuse: malware URLs, rate limits?
- 07301 vs 302 — caching implications?
- 08Multi-region redirects day one?
Capacity → decisions (label numbers as heuristics)
1# Scenario A — moderate product (interview default)2new_links_per_day = 100_000_000 / 365 # ~3e5 writes/day if 100M/year3write_qps = 3e5 / 86400 # ~3–4 QPS avg; peak ~20 QPS4redirects_per_day = 50_000_000 # example5read_qps = redirects_per_day / 86400 # ~600 avg → peak ~2–3K6# ID space: 7-char base62 = 62^7 ≈ 3.5e12 — huge vs 1e8 links/year7# Storage: ~200B/row × 1e8/year ≈ tens of GB/year raw — single DB fine early8# Decision: ID space is not the bottleneck; read QPS + cache is.910# Scenario B — bit.ly-class (secondary-cited / heuristic — label it)11# ~10B redirects/day → ~115k RPS avg → ~500k+ peak with spike factor12# 99% cache hit → origin sees ~1% → still large; edge cache mandatory13# 1B links × 200B = 200GB working set class → multi-node Redis/Memcached14# Write 1B links/day → ~11.5k write QPS avg → sharded metadata store1516# Numbers → decisions: analytics never on 302 path; shard when write QPS17# exceeds one primary; CDN/edge for viral codes.Key idea
API & data model
1POST /api/v1/links2 Headers: Idempotency-Key, Authorization?3 Body: { long_url, custom_alias?, expires_at? }4 → 201 { code, short_url }56GET /{code} → 302 Location: long_url (or 301 if permanent + cacheable intentionally)7DELETE /api/v1/links/{code}8GET /api/v1/links/{code}/stats → aggregates (async path)910links(11 code PK, -- base62 or custom12 long_url,13 owner_id NULL,14 created_at,15 expires_at NULL,16 is_custom BOOL17)18clicks(short_code, ts, ip_hash, ua, country, referer) -- partition by day/code; async only1ALPHABET = "0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz"23def base62_encode(n: int, width: int = 7) -> str:4 out = []5 while n:6 n, r = divmod(n, 62)7 out.append(ALPHABET[r])8 s = "".join(reversed(out)) or "0"9 return s.rjust(width, "0")[-width:]1011# Counter service hands out [start, end) ranges per instance — no central lock per request.1213ID STRATEGY TRADEOFF14Approach Pros Cons Pick when15-------------------- --------------------------- --------------------- --------------------16MD5/SHA truncate simple collisions, retries internal tooling17Base62 counter monotonic, compact hot row if unsharded low-write sortable18Snowflake distributed, time order NTP/config care multi-DC high QPS19UUID v7 no coord, sortable longer keys modern internal IDs20Random base62+unique simple ops retries under load many interviewsCommon mistake
“UUIDs make great short links.”
Common mistake
“We need Cassandra on day one because shortener interviews always use it.”
Key idea
1def redirect(code: str):2 row = cache.get(code)3 if row is None:4 row = db.get_link(code) # primary or read replica OK if create→redirect lag rare5 if row is None:6 return 4047 cache.set(code, row, ttl=3600)8 if row.expired():9 return 41010 enqueue_click(code) # async; never block on analytics DB11 return 302, row.long_urllogin, admin); consider separate table or flag is_custom for policy. Celebrity tweet of your short link is a first-class load test.Analytics path
UPDATE click_count on the hot row inline for viral codes — that creates a write hotspot on the most popular keys.Scale-up story & failure modes
Senior close: “Redirect path is cache + thin service; writes allocate IDs from ranges; analytics is fully async; hot campaign codes are a first-class case.”
Interview answers — URL shortener
- 01Key length? → 7 base62 ≈ 3.5T; years of headroom at 1M creates/day class.
- 02Hash collisions? → unique constraint + retry, or Snowflake/counter to avoid.
- 03Globally fast redirects? → geo replicas + edge/CDN cache.
- 04Abuse? → CAPTCHA, rate limit, IP reputation, malware scan async.
- 05Custom aliases break? → uniqueness, reserved words, policy path.
- 06Expiry? → TTL cache + DB sweep/tombstone.
- 07Analytics without slowing 302? → fire-and-forget queue.
- 08Celebrity tweet? → pre-warm cache/CDN; protect origin.
- 09Why not UUID in the URL? → too long; fine as internal id.
- 10Scale writes 100k/s? → range IDs, shard metadata, batch inserts.
- 11301 vs 302? → 301 caches hard at clients; 302 keeps control.
- 12Idempotent create? → Idempotency-Key returns same code.
Design a URL Shortener / TinyURL (search Gaurav Sen)Gaurav SenarticleHelloInterview — Design Bitly / URL shortenerHello InterviewarticleByteByteGo — Design a URL ShortenerByteByteGoarticleSnowflake ID announcement (historical)Twitter Engineering (historical)docsStripe Idempotent requestsStripedocsMDN HTTP 302MDN301 vs 302 and CDN caching (deep)
1REDIRECT CACHING MATRIX2Layer What to cache TTL mindset3--------------- ------------------------- --------------------------4Browser optional for 301 long if permanent5CDN edge public non-auth codes short (30–300s) + purge API6App Redis code → long_url minutes–hours; delete on edit7Local process ultra-hot codes seconds; single-flight fill89Never cache personalized analytics pages the same way as 302 Location.Abuse, malware, and reserved aliases
login, admin, brand names), and a disable flag that edge can honor from a push config. Senior answers treat abuse as a design surface, not an afterthought bullet. Analytics productization: minute rollups for dashboards, raw events in cheap object storage for ad-hoc, privacy (hash IPs, careful with PII in query strings of long URLs). If the interviewer asks for “real-time counters,” offer approximate (Redis HyperLogLog / rolling counters) on the hot path and exact batch later.Multi-region shortener evolution
1# ID space with region nibble (sketch)2# [timestamp | region | worker | sequence] → base623# Guarantees no cross-region counter collision without a global lock4# Custom aliases still need a global uniqueness check (harder problem)Checkpoint
Where should the cache live for redirects, and why?
Checkpoint
Two create requests with the same Idempotency-Key arrive. Correct behavior?
Checkpoint
Why pre-allocate Snowflake/counter ranges instead of hitting a global sequence on every write?
Checkpoint
A celebrity posts one short link; redirect QPS spikes 100×. What do you cut/change first?
Checkpoint
Why might estimation show a single DB is enough for metadata yet you still introduce Redis?
End-to-end narration script (12 minutes)
1SHORTENER DECISION LOG (write on board)2Assumption Number (heuristic) Decision3--------------------- --------------------- --------------------------4Creates/day 3e5 single primary OK early5Redirects/day 5e7 cache-aside mandatory6Peak redirect RPS ~3k (×5) few app nodes + Redis7Key length base62×7 decades of headroom8Viral code share 10–50% of QPS CDN + local cache9Analytics volume = redirects async only; never 302 path10Custom alias rate low but abusive policy + unique + reserved1112Rejected: UUID short links; sync malware scan on redirect; global lock IDs.URL shortener math + decision card
1MODERATE PRODUCT (heuristic)2creates/day 3e5 → write QPS ~3–4 avg, ~20 peak3redirects/day 5e7 → read QPS ~600 avg, ~3k peak4row size ~200B → storage tens of GB/year — single DB OK early5base62 len 7 → 62^7 ≈ 3.5e12 codes — space not the bottleneck6THEREFORE: cache redirects; keep metadata SQL/primary simple; async clicks78BITLY-CLASS ESCALATION (secondary/heuristic — label it)9redirects/day 1e10 → ~1e5 RPS avg → peak much higher10cache hit 99% still leaves large origin QPS without edge11metadata shards + multi-node cache + CDN required12analytics warehouse path mandatory1314ID PICKER15low write + simple → random base62 + unique retry16high write multi-DC → Snowflake / range counters17custom alias → separate uniqueness + reserved words1819HOT PATH RULES201) 302 decision uses cache/DB only212) enqueue click never blocks response223) single-flight on hot miss234) disable flag for malware must reach edge fast245) 301 only if permanent + purge story exists2526FAILURE DRILLS27cache stampede, counter service down, alias race 409,28analytics lag, DB outage serve hot from cache, abuse floodCould you drive a full shortener design: requirements, capacity→ID/storage choice, write/read paths, hot key, async analytics?
Takeaways
- Estimation either kills over-engineering or forces edge cache — depending on numbers.
- Base62 + counter ranges or random+unique; customs are a policy path.
- 302 path: cache-aside, expiry, async clicks; never inline viral counters.
- Idempotent creates; unique codes; hot alias plan.
- Label bit.ly-class figures as reported/heuristic when you cite them.
Next: newsfeed at scale — fan-out-on-write vs pull, celebrities, ranking, and cost per render.
Sources
Free to read · better with Enzo
Learn it with Enzo
Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.