Marketplace & Booking
Design URL Shortener
01
Requirements
Requirements
- Users can shorten any URL to a 7-character code
- Anyone with the short link gets redirected to the original URL
- Links can optionally have a custom expiry date
- Creators can view click analytics — count, geography, device
- Users can optionally provide a custom alias instead of auto-generated code
- 99.99% availability — broken links are catastrophic
- Redirects under 100ms p99 globally
- Short codes must be globally unique — zero collision tolerance
- Read-heavy: 1000:1 read-to-write ratio
- Eventual consistency acceptable — 50ms replication lag is fine
Before designing, always ask: Do links expire or live forever? Do you need click analytics? Are custom aliases required? Expected scale — 1M DAU or 100M DAU? These four questions change the architecture significantly.
02
Scale Estimation
Scale Estimation
| Metric | Calculation | Result |
|---|---|---|
| Daily Active Users | bit.ly scale assumption | 100M |
| Daily writes (new links) | 100M × 1% create links | 1M / day |
| Write RPS (average) | 1,000,000 ÷ 86,400 | ~12 / sec |
| Write RPS (peak) | 12 × 5× peak factor | ~60 / sec |
| Daily reads (clicks) | 100M × 10 clicks/user | 1B / day |
| Read RPS (average) | 1,000,000,000 ÷ 86,400 | ~11,600 / sec |
| Read RPS (peak) | viral moments | ~100,000 / sec |
| Read : Write ratio | 11,600 ÷ 12 | ~1,000 : 1 |
| Storage (5 years) | 1M/day × 365 × 5 × 500B | ~1 TB |
| Short code space (7 chars) | 62^7 | 3.5 trillion |
| Runway at current write rate | 3.5T ÷ 1M/day | ~9,500 years |
The 1,000:1 read-to-write ratio is the single most important number in this system. It tells us writes are trivially easy (PostgreSQL handles 60/sec in its sleep) and reads are the entire engineering challenge. Every architectural decision — Redis, CDN, LFU eviction, hotness tiers — exists because of this ratio.
Storage is not a bottleneck. 1TB over 5 years fits on a single modern SSD. This tells you the problem is throughput, not capacity. You are not building a data warehouse — you are building a high-speed lookup service.
03
API Design
API Design
Returns 204 No Content. Soft-delete — record purged after 30 days. Must also DEL from Redis and purge CDN entry synchronously.
Always use 302 Temporary Redirect, never 301. A 301 tells browsers to cache the redirect permanently — after the first click, your servers never see subsequent clicks. You lose all analytics. Since analytics is the commercial value of a URL shortener, using 301 destroys your product while appearing to optimise it.
04
High-Level Architecture
High-Level Architecture
Tier 2 and Tier 3 viral links live here. A user in Tokyo clicking a viral link never touches your origin servers — the CDN edge node responds in under 5ms. The hotness monitor pushes links to CDN when they cross 1,000 clicks/minute.
Stateless compute — any instance handles any request. This enables horizontal scaling. Session state lives in Redis, not in memory. Each server claims keys from the pool, never generates them on the fly.
LFU eviction protects historically popular links. TTL jitter (±1hr) prevents expiry cliffs. Sliding window counters track clicks/minute per link. A cache hit costs <1ms vs ~5ms for a DB read.
Pre-generated 7-char Base62 codes stored in an keys_available table. FOR UPDATE SKIP LOCKED prevents race conditions across servers. The background generator maintains 10M+ codes available at all times.
05
Deep Dive — Key Generation & Hotness Tiers
Deep Dive — Key Generation & Hotness Tiers
This problem has two technically interesting cores: how you generate unique short codes safely across distributed servers, and how you detect and serve viral links without touching your database. Everything else is standard web infrastructure.
sequenceDiagram
participant C as Client
participant CDN as CDN Edge
participant API as API Server
participant R as Redis
participant DB as PostgreSQL
C->>CDN: GET /x9kZ3mP
alt Tier 2/3 — CDN hit
CDN-->>C: 302 redirect (< 5ms globally)
else CDN miss
CDN->>API: Forward request
API->>R: GET x9kZ3mP
alt Redis hit
R-->>API: long_url (< 1ms)
API-->>C: 302 redirect (~10ms total)
else Redis miss
API->>DB: SELECT long_url WHERE short_code = 'x9kZ3mP'
DB-->>API: long_url (~5ms)
API->>R: SET x9kZ3mP TTL=86400 (async)
API-->>C: 302 redirect (~20ms total)
end
end
Key Generation — The Pre-Generated Pool
Three approaches exist for generating short codes. Random on-demand requires a DB read on every write to check for collisions. MD5 hashing is deterministic (deduplication for free) but needs salt-and-rehash on collision. The pre-generated key pool wins because all uniqueness checking happens offline in the background — the write path just claims a key atomically.
The critical SQL clause is FOR UPDATE SKIP LOCKED. When 10 API servers simultaneously claim keys, each one locks a row and any server that finds its target already locked simply skips to the next available key. No race condition. No collision. No coordination overhead. The entire operation is a single atomic transaction.
Hotness Monitor — Three-Tier Traffic Detection
Every click increments a Redis counter with minute-level granularity: INCR clicks:x9kZ3mP:2024-01-15-14:37. The hotness monitor sums the last 5 minute-buckets every 60 seconds to get a stable clicks/minute figure for each active link. Links are then sorted into three tiers based on traffic thresholds.
Tier 1 (100–999/min) — Redis cache with TTL refreshed on access. Tier 2 (1,000–9,999/min) — Redis plus CDN edge globally, TTL renewed every 60s by the monitor. Tier 3 (10,000+/min) — pre-computed static redirect served at CDN with no application logic involved at all. Demotion is passive — when traffic drops, the monitor stops renewing the CDN TTL, and it expires naturally within 5 minutes.
Why 7 Characters? The Math.
7 characters from Base62 (a–z, A–Z, 0–9) gives 62^7 = 3.5 trillion unique combinations. At 1 million new links per day, that's 9,500 years of runway. 6 characters would give 56 billion combinations — enough mathematically but without safety margin. 8 characters is unnecessary. 7 is the sweet spot derived from requirements, not guesswork.
The avalanche effect in Base62-encoded hashes ensures inputs are uniformly distributed across the entire 3.5 trillion slot space. Even after 2 billion stored URLs, you're using only 0.057% of available space — collision probability per new insert is approximately 0.029%.
06
Key Design Decisions & Tradeoffs
Key Design Decisions & Tradeoffs
All uniqueness work happens offline. Write path just claims a key atomically with FOR UPDATE SKIP LOCKED. Zero collision possible. No DB reads on write path. Background generator easily keeps pace with 60 writes/sec.
Random requires a DB existence check on every write (extra round trip). MD5 hashing gives free deduplication but requires salt-and-rehash on collision. Both work at small scale but add reactive complexity vs the pool's proactive approach.
~ Acceptable at low scaleEvery click hits your servers. Full analytics on every request — count, device, country, referrer. The short URL can be updated or expired at any time. Analytics is the product — this is the only viable choice.
✓ Required for analytics-driven business modelBrowser caches redirect forever after first click. Near-zero server load for returning users. But you lose all subsequent click data. You cannot update the destination. You cannot expire the link. Destroys your product while appearing to optimise it.
✗ Kills analytics — only for internal toolsEvicts by total access frequency. A link with 10M historical clicks that went quiet 2 hours ago stays protected. URL access follows a power law — historically popular links predict future clicks better than recent ones.
✓ Right for power-law access patternsEvicts by recency. Redis default. Works well for most systems but can evict a viral link during a brief quiet period, causing a thundering herd on the next traffic spike. Suboptimal for URL shorteners specifically.
~ Fine for generic caching, suboptimal hereACID guarantees, FOR UPDATE SKIP LOCKED for key pool, complex queries, familiar tooling. Handles 60 writes/sec using 1.2% of capacity. Read replicas scale reads to 100x. Simple schema, proven at this scale.
Massive write throughput, linear horizontal scaling, no single point of failure. But gives up ACID, joins, and strong consistency. Adds significant operational complexity for a problem that doesn't need it. PostgreSQL handles our write load in its sleep.
~ At 1000× scale only07
What Can Go Wrong
What Can Go Wrong
A viral cached link expires. Thousands of requests simultaneously miss the cache and all query the database at once. DB CPU spikes, latency climbs, cascading failure possible.
→ Fix: Mutex lock on cache miss — only one request queries DB, rest wait for cache populationCelebrity posts a brand new link. It has never been cached. 100,000 simultaneous first-clicks all miss cache at the same moment before the hotness monitor can promote it.
→ Fix: Hotness monitor promotes to CDN within 60s. Redis absorbs the burst in the interim — it handles 100K ops/sec comfortably.Primary DB goes down. All writes fail — new links cannot be created. Reads survive on replicas but may serve slightly stale data during the failover window.
→ Fix: AWS RDS Multi-AZ automated failover. Standby promoted in ~30s. Reads unaffected — replicas still serve. CDN protects hot links entirely.Cache goes dark. 100% of reads fall to the database immediately. At 11,600 reads/second, a single PostgreSQL instance is overwhelmed within seconds — not minutes.
→ Fix: Redis Sentinel for automatic failover (<60s). CDN provides partial safety net — Tier 2/3 links continue serving from edge nodes unaffected.Background key generator crashes silently. Pool drains at 60 writes/sec. When pool hits zero, write requests fail — users cannot shorten new URLs.
→ Fix: Alert at <1M keys remaining (46 hours runway). Fallback to random generation with collision check. Generator restart resolves within minutes.User updates their short link destination. DB is updated immediately but Redis and CDN still serve the old URL. Users get redirected to the wrong place — a correctness bug, not just a performance issue.
→ Fix: On every URL update, explicitly DEL from Redis and call CDN purge API synchronously before returning success to the user.Entire datacenter or cloud region goes dark. All links hosted in that region become inaccessible. Every QR code, marketing email, and printed poster pointing to those links breaks simultaneously.
→ Fix: Multi-region active-active deployment. DNS failover routes to healthy region in <60s. CDN edge nodes are globally distributed — hot links unaffected by origin region failures.08
Interview Tips
Interview Tips
Clarify before drawing anything. Ask four questions: Do links expire? Do you need analytics? Custom aliases? Expected scale? These change the architecture significantly. Interviewers reward candidates who treat requirements as inputs, not assumptions.
Lead with the read/write ratio. Say it explicitly: "This system is 1,000:1 read-heavy — that single number defines the architecture." Everything that follows — Redis, CDN, LFU — should trace back to this ratio. Interviewers want to see you derive architecture from data.
Avoid the NoSQL trap. Most candidates jump to Cassandra or DynamoDB to "handle scale." This signals poor judgment. PostgreSQL handles 60 writes/sec using 1.2% of capacity. Say: "I'd start with PostgreSQL and migrate to NoSQL only if I hit its write ceiling" — which won't happen at this scale.
Name all three key generation options, then justify the pool. Random → MD5 hash → pre-generated pool. The magic phrase: "FOR UPDATE SKIP LOCKED prevents race conditions atomically across distributed servers." Knowing this SQL clause signals genuine production database experience.
The 301 vs 302 question always comes up. Answer: "301 destroys analytics — the entire commercial value of a URL shortener is click tracking, so 302 is the only viable choice." Candidates who say 301 don't understand the business model.
The hotness monitor is your differentiator. Most candidates stop at "put Redis in front of the DB." Describing a sliding window counter that detects tier transitions and pushes viral links to CDN edge separates good answers from great ones. It shows you think in systems, not just components.
Name your failure modes before being asked. Say proactively: "The thundering herd is the main risk — a mutex lock on cache miss means only one request queries the DB while others wait." Candidates who raise problems before being asked look senior. Those who wait to be asked look reactive.
End with the evolution story. "I'd start with a monolith — one server, one PostgreSQL, no Redis. Ship fast." Then walk through: add Redis → add read replicas → add CDN tiering → shard DB → multi-region. Each step triggered by a specific metric. This shows engineering maturity — solutions proportional to problems.
09
How the Design Evolves
How the Design Evolves
Each phase is triggered by a specific metric crossing a threshold — not by time, not by team size, not by intuition. Premature optimisation is the enemy. Every phase adds complexity that must be justified by a real problem.
One server, one PostgreSQL database, no cache, no queue. Random code generation with collision check is fine at this scale. Ship fast. This handles ~500 reads/sec easily. Most startups live here for years. The correct answer for MVP is boring infrastructure.
DB read latency starts climbing. Add Redis cache with LFU eviction in front of PostgreSQL. Migrate to the pre-generated key pool to eliminate write-path collision checks. Add a read replica. Most products live here their entire lifecycle. This configuration handles ~50K reads/sec.
Cache hit rate is high but viral spikes still stress the origin. Add CDN edge caching for hot links, driven by the hotness monitor. Add more read replicas. API servers scale horizontally behind the load balancer. Redis cluster for cache HA. This is the architecture described in this document — the "full" design.
PostgreSQL write throughput approaches its ceiling (even at 60 writes/sec, data volume and index size create operational complexity at this scale). Shard by short_code hash range across multiple primaries. Deploy full active-active stack in multiple regions. Global load balancing with GeoDNS. Most engineers never operate here — but knowing it demonstrates architectural depth.
At true planet scale (think bit.ly or t.co), consider migrating the hot path entirely to a NoSQL store like Cassandra for the short_code → long_url lookup table. Custom hardware at CDN PoPs. Dedicated analytics pipeline (Kafka → Flink → data warehouse). The read path becomes a pure key-value lookup with no application logic — raw bytes from edge nodes in under 2ms globally.
Watch and read
References & Videos
Try next
Free to read · better with Enzo
Whiteboard this with Enzo
Enzo runs it as a live system design round on the whiteboard and grades your trade-offs.