Lesson 6 of 8 · 55 min
Worked design — realtime chat & notifications
WebSocket gateways, presence, per-channel ordering, multi-device catch-up, push reliability, and message persistence with Discord’s reported production anchors.
Lesson 6 · Worked design
Realtime chat & notifications
Connections are state
Capacity sketch (decision-linked)
1# Heuristic capacity — state assumptions2concurrent_ws = 50_000_0003per_gateway = 50_000 # heuristic conns/box4gateways = concurrent_ws / per_gateway # ~1000 gateway boxes class56msgs_day = 100_000_000_000 # 100B/day class product (illustrative)7msg_qps_avg = msgs_day / 86_400 # ~1.2M msg/s avg8msg_qps_peak = msg_qps_avg * 8 # large peak factor9bytes_msg = 50010# storage 5y × 3 replicas → multi-hundred-PB class cold → MUST tier cold to object store1112# DECISIONS FORCED:13# - gateway fleet + sticky/resume sessions14# - partition message log by channel_id (+ time bucket)15# - pub/sub for live fan-out separate from durable store16# - cold tier for history older than weeks/months17# - push path separate with coalesce/throttleHigh-level topology
1Clients (mobile/web)2 │ WebSocket3 ▼4Gateway fleet ── session: user/device → conn · heartbeats · reconnect resume5 │6 ├─► Channel/router service (who is in channel? which gateways?)7 ├─► Pub/sub / Kafka (channel partition key)8 ├─► Message service → durable store (Cassandra/Scylla-class or equiv)9 ├─► Presence service (often Redis / ephemeral)10 └─► Push worker → FCM/APNs (offline devices)1112History API (HTTP) for catch-up / infinite scrollchannel_id; single-threaded virtual consumer per partition; assign monotonic channel_seq on assign. Clients may still see retries — use server seq + client nonce for idempotent send. Discord-style buckets: partition key (channel_id, time_bucket) so a single channel’s history does not create one unbounded partition forever.1def send_message(channel_id, sender_id, body, client_nonce):2 if dedupe.exists(channel_id, sender_id, client_nonce):3 return dedupe.get(...) # retry-safe4 seq = next_seq(channel_id) # from partition owner / atomic counter5 msg = Message(id=new_id(), channel_id, seq, sender_id, body, ts=now())6 store.append(msg) # durable first or after assign — state choice7 pubsub.publish(channel_id, msg)8 dedupe.put(...)9 return msgKey idea
Key idea
Presence, typing, receipts
Multi-device & offline
Common mistake
“Push notifications guarantee message delivery.”
Persistence — Discord-reported lessons
1# Message storage sketch (channel-bucketed)2# PK: (channel_id, bucket, msg_seq) bucket = e.g. yyyyww or snowflake range3# Access: latest page = max bucket for channel; scroll up = previous bucket45STORAGE CHOICE TRADEOFF6Store Pros Cons Pick when7-------------------- ----------------------------- ----------------------- -------------------8Cassandra/Scylla write scale, PK lookups ops/compaction care high-write chat9Postgres + replicas familiar, joins, tx write ceiling small team chat10Kafka→S3 + search replay + audit cold read latency regulated history1112# Hot channel mitigation: separate outbox/fanout tier for large member listsNotifications fan-out & 10× failures
Common mistake
“We need total global ordering of all messages in the product.”
Senior close: “Per-channel order, durable append store, WS gateway + pub/sub fan-out, catch-up by seq, push as best-effort, presence as ephemeral.”
Interview answers — chat
- 01Exactly-once delivery? → you can’t; at-least-once + client/server dedupe by msg id/nonce.
- 02Join channel with 1M history? → don’t backfill all; recent + lazy scroll.
- 03Shard by channel not user? → per-channel order needs single writer/partition.
- 04Typing without flood? → debounce + short TTL pub/sub.
- 05Group read receipts? → per-user counters, not message row rewrites.
- 06Offline delivery? → durable store + push wake + catch-up by last_seq.
- 07Search messages? → async inverted index (L7), not OLTP LIKE.
- 08E2E encryption? → client keys; server metadata limited; different design space.
- 09Scale WS gateways? → shard conns, resume tokens, Envoy/L7 routing class.
- 10Admin broadcast to 10M? → async fanout workers + badge batching.
- 11Why cite Discord migration? → reported anchor for append-heavy storage + p99/ops.
- 12Hot channel? → time buckets + fanout tier; throttle degraded mode.
Message API and history pagination
1WS /ws — auth, subscribe channels, send, ack, resume(last_seq)2POST /channels/{id}/messages — REST fallback3GET /channels/{id}/messages?cursor&limit=504POST /channels/{id}/typing56messages PK idea: (channel_id, bucket, seq)7memberships: (user_id, channel_id, last_read_seq)8cursor = (bucket, seq) — scroll upward loads older bucketLarge channels and fan-out amplification
1CHAT SENIOR SIGNALS2MID FAILURE SENIOR SIGNAL3-------------------------------------- ------------------------------------------4messages are just rows per-channel order is first-class5one fanout approach for DM and 10k room different delivery strategies6mix history store with live pubsub separate tiers, shared message id/seq7push = delivery push = wake; store = truth8ignore E2EE when asked E2EE changes metadata visibility9just add a column on 1T messages migration/backfill strategyReconnect, resume, and catch-up storms
Checkpoint
What ordering guarantee should you claim by default for group chat?
Checkpoint
User has phone offline and laptop online. Message arrives. Correct path?
Checkpoint
Why mention Discord’s Cassandra→Scylla migration in an interview?
Checkpoint
Send API is retried after a gateway blip. How do you avoid duplicate messages in-channel?
Checkpoint
QPS jumps 10× after a viral event. First healthy degradation?
End-to-end narration script (15 minutes)
1CHAT DECISION LOG2Signal Forces3-------------------------- ------------------------------------4Per-channel order partition by channel_id (+ bucket)5At-least-once networks client nonce dedupe6Multi-device durable store + per-device push/WS7Offline push best-effort + catch-up by seq8Long retention hot/cold tiers; object store cold9Large rooms delivery fanout tier ≠ storage tier10Outage recovery jittered resume; rate-limit catch-up1112Reported anchor: Discord Cassandra→Scylla ~177→72 nodes; p99 improved (2023 blog).Chat math + decision card
1HEURISTIC WORKED EXAMPLE250M concurrent WS / 50k per gateway → ~1000 gateways3100B msgs/day → ~1.2M msg/s avg (illustrative product scale)4500B bytes/msg → multi-PB retention with RF → cold tier required5THEREFORE: gateway fleet, channel partitions, hot/cold storage, push separate67REPORTED ANCHORS (Discord 2023 blog — not our metrics)8Cassandra ~12 nodes (2017) → ~177 nodes (trillions msgs early 2022)9Scylla migration → ~72 nodes; p99 read ~15ms / write ~5ms class steady10partition key (channel_id, time_bucket) bounds hot channels1112GUARANTEES13order: per-channel not global14delivery: at-least-once + nonce dedupe15push: best-effort wake; store is truth16presence/typing: ephemeral TTLs1718SEND PATH19dedupe nonce → assign seq → durable append → pubsub → gateways20REST history for catch-up with cursor (bucket, seq)212210x DEGRADE ORDER23scale gateways → coalesce push → rate-limit catch-up/presence24→ backpressure sends → NEVER drop durability first2526LARGE ROOM27storage still one stream; delivery fanout is separate amplification problemCan you design chat with WS gateway, per-channel order, multi-device catch-up, and a durable store story with cited production anchors?
Takeaways
- Gateway + pub/sub + durable message service is the spine.
- Order per channel via partitions/seq; idempotent sends with nonces.
- Presence is ephemeral; push is best-effort; store is truth.
- Discord blog numbers are reported anchors for storage/latency lessons.
- 10×: scale gateways, backpressure, coalesce notify — don’t drop durability.
Next: two high-frequency component designs — distributed rate limiting and search/inverted index.
Sources
Free to read · better with Enzo
Learn it with Enzo
Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.