Lesson 6 of 8 · 58 min

FE system design I — chat & realtime dashboards

Full RADIO frontend system-design skeletons for chat (optimistic send, virtualization, scroll anchoring, reconnect) and realtime dashboards (SSE/WS, downsampling, backpressure) with interview timing and NFRs.

Lesson 6 · FE system design I

Chat & realtime dashboards

Drive the whiteboard

Senior FE loops almost always include a frontend system design round. You clarify product requirements, lock NFRs, pick transport and state, design rendering for scale, then deep-dive failure modes. This lesson walks two full RADIO skeletons: chat and realtime dashboards.
Framework (RADIO, FE-flavoured): Requirements · Architecture (clients, transport, storage) · Data model · Interfaces (API/events) · Observability & risks. Always name a11y and perf budgets before the interviewer forces them. Stay FE-primary; sketch backend only as needed.

How to run the 45 minutes

text
10–5m   Clarify: users, platforms, scale, offline?, multi-device?, SEO?25–10m  NFRs: latency, delivery guarantees, a11y, perf budgets310–25m High-level: surfaces, state buckets, transport, server roles425–40m Deep dives: the hard parts (scroll, races, backpressure…)540–45m Risks, metrics, rollout, what you'd build first67Trap: spending 30m on DB schema. Stay FE-primary.

Prompt A — Chat / messaging UI

Product: 1:1 and small group chat, mobile+desktop web, near-realtime delivery, typing indicators, read receipts, offline compose, multi-device. Scale: millions of users, hot conversations with high message rate.

Chat — full RADIO skeleton

text
1R — REQUIREMENTS & NFRs2  Functional: send/receive, history scroll-up, typing, read receipts, media3  NFR: send feels <100ms optimistic; eventually consistent multi-device;4       10k+ msgs/thread virtualized; a11y keyboard + live regions without spam;5       INP on composer; CLS when prepending history; offline outbox6  Out of scope (ask): E2E crypto, voice/video, bots platform78A — ARCHITECTURE9  ┌──────── client ────────┐   ┌──── edge/API ────┐   ┌── services ──┐10  │ Composer (local draft) │──▶│ REST/Action send │──▶│ msg service  │11  │ Virtualized msg list   │◀──│ WebSocket fanout │◀──│ presence     │12  │ Connection manager     │◀──│ resume/catch-up  │   │ media        │13  │ Offline outbox (IDB)   │   └─────────────────┘   └──────────────┘14  State buckets:15    server cache: message pages by channel+cursor16    URL: /c/:channelId (+ thread id)17    local: draft, scroll pin, emoji picker18    global: connection toast, unread map in shell1920D — DATA MODEL (client)21  Msg { id, clientId, channelId, body, status: pending|sent|failed,22        createdAt, seq?, senderId }23  Page { channelId, cursor, items[] }  // cursor not offset24  Presence { typingUserIds[], lastReadByUser }2526I — INTERFACES27  HTTP: POST /messages { clientId, body } → Msg28        GET /messages?channel&cursor&limit → { items, nextCursor }29  WS events: message.created | message.ack | typing | read | presence30  Idempotency: clientId on send; server de-dupes3132O — OBSERVABILITY & RISKS33  Metrics: ws uptime, catch-up lag, send fail rate, INP composer,34           dropped frames, time-to-visible on open channel35  Risks: reconnect storm, out-of-order, scroll jump, multi-tab sockets36  Rollout: flag realtime v2; chaos kill WS in QA
Transport: WebSocket (or HTTP/2 SSE for one-way) for fanout; HTTP for history pages and media upload. SSE is simpler through some proxies but weaker for bidirectional typing/presence — call the tradeoff. At-least-once delivery + client idempotency keys; order via monotonic seq per channel.

Chat — hard FE problems (deep dives)

  1. 01Optimistic send + rollback — temp client id → replace with server id; mark failed; retry from outbox; never silent delete.
  2. 02Scroll anchoring — prepending history must not jump the viewport; use anchors / scrollTop compensation. Most common chat FE bug.
  3. 03Virtualization — dynamic row heights (media) need measured size cache (Virtuoso / TanStack Virtual).
  4. 04Presence & typing — throttle events; ephemeral state via presence service; do not persist typos forever.
  5. 05Read receipts state machine — sent → delivered → read; multi-device last-read merge.
  6. 06Multi-device — conflict when edits/deletes race; last-write or CRDT for edits if required.
  7. 07Reconnect — jittered exponential backoff; resume token / last event id; catch-up query; dedupe by message id.
  8. 08Multi-tab — BroadcastChannel or navigator.locks so one tab owns the socket.
tsx
1// Optimistic send sketch2type Msg = {3	id: string4	clientId: string5	status: 'pending' | 'sent' | 'failed'6	body: string7	createdAt: number8}910async function send(body: string) {11	const clientId = crypto.randomUUID()12	const optimistic: Msg = { id: clientId, clientId, status: 'pending', body, createdAt: Date.now() }13	appendLocal(optimistic)14	try {15		const saved = await api.send({ clientId, body })16		replaceByClientId(clientId, { ...saved, status: 'sent' })17	} catch {18		markFailed(clientId) // retry UI + outbox19	}20}2122// Reconnect with jitter23const delay = Math.min(30_000, 500 * 2 ** attempt) + Math.random() * 250

Prompt B — Realtime analytics dashboard

Product: live ops dashboard — time-series charts, alert table, filterable dimensions, multi-widget layout. Updates 1Hz+. Multiple users. Accessibility for dense data. Dark mode.

Dashboard — full RADIO skeleton

text
1R — REQUIREMENTS & NFRs2  Live mode 1Hz+ ticks; historical range query; filter dimensions;3  multi-widget layout; alert table 50k rows; keyboard alternatives to charts;4  colourblind-safe series; p75 INP ok while streaming; per-widget error states56A — ARCHITECTURE7  Data path options:8    A) SSE stream of tick aggregates from query service9    B) WebSocket multiplexed channels per widget10    C) Polling 5–15s when live mode off (cheaper)11  Client pipeline:12    network → normalize → downsample → ring buffer → chart lib13                   ↘ backpressure drop/coalesce when UI busy14  State: URL time range/filters/selected incident;15         server historical pages; local live toggle/hover; global theme/auth1617D — DATA MODEL18  Tick { metricId, ts, value, dims }19  SeriesRingBuffer { maxPoints, points[] }  // pixel-budgeted20  AlertRow { id, severity, openedAt, ... }2122I — INTERFACES23  GET /metrics/query?range&dims → historical24  SSE /metrics/stream?subs=... → tick events25  GET /alerts?cursor= → virtualized table pages2627O — OBSERVABILITY & RISKS28  Metrics: dropped ticks, widget render ms, stream lag, error rate per widget29  Risks: backpressure melt, memory growth, colour-only encoding, full-page spinner

Dashboard — hard FE problems

  1. 01Backpressure — coalesce frames (rAF), drop intermediate ticks, or sample.
  2. 02Downsampling — render max points for pixel width (e.g. LTTB); keep raw in worker if needed.
  3. 03Virtualize tables — 50k alert rows never full-DOM.
  4. 04Color/contrast — dense series; colourblind-safe palettes; do not encode only by colour.
  5. 05Keyboard — chart alternatives (data table), focusable legend toggles.
  6. 06Empty/error/loading — per widget, not one full-page spinner.
ts
1// Backpressure with rAF coalesce2let latest: Tick | null = null3let scheduled = false45socket.on('tick', (t) => {6	latest = t // keep only newest if overwhelmed7	if (!scheduled) {8		scheduled = true9		requestAnimationFrame(() => {10			scheduled = false11			if (latest) applyTick(latest)12		})13	}14})

Shared comparison: WS vs SSE vs poll

  1. 01WebSocket — bidirectional, presence/typing; reconnect complexity; proxy issues.
  2. 02SSE — server→client, auto-reconnect in EventSource; simple for feeds/ticks; no native binary client→server.
  3. 03Poll — simplest ops story; higher latency/load; fine for soft realtime.

Senior signals for this round

  1. 01Distinguishes fan-out (push) from fan-in (write path).
  2. 02Names backpressure: max buffer + per-frame budget.
  3. 03Justifies WS vs SSE with direction and infra constraints.
  4. 04Models failure: network drop, tab close, slow consumer, reconnect storm.
  5. 05Budgets a11y + perf unprompted in the first ten minutes.
System Design: Chat Application — Hussein NasserHussein Nasser

Chat RADIO — deep dive scripts (say these out loud)

Optimistic send narrative: User hits Enter. Client appends a message with clientId and status pending. UI scrolls if pinned to bottom. Network POST (or WS) carries clientId for idempotency. On ack, replace temp id with server id and set status sent. On failure, mark failed with retry; outbox persists across reload via IndexedDB if offline is in scope. Never silent-drop — trust is the product.
Scroll anchoring narrative: User is reading older messages. IntersectionObserver near the top fires. Client fetches next older page by cursor. Before DOM insert, record scrollHeight and scrollTop. After prepend, set scrollTop = oldScrollTop + (newScrollHeight - oldScrollHeight). Dynamic media rows need a size cache so virtualization does not thrash. Name this unprompted — interviewers wait for it.
Reconnect narrative: Socket closes. Show reconnecting status. Backoff with jitter: min(30s, 500ms * 2^attempt) + random(0..250ms). On open, send lastSeq or resume token; server replays missed events; client de-dupes by message id. Multi-tab: elect a leader with navigator.locks or BroadcastChannel so five tabs do not open five sockets.
ts
1// Presence typing throttle2let lastSent = 03function onKeystroke() {4	const now = Date.now()5	if (now - lastSent < 300) return6	lastSent = now7	ws.send(JSON.stringify({ type: 'typing', channelId }))8}9// Ephemeral: server fans out; clients clear typing after ~3s without events10// Do not write typing rows to the messages table

Dashboard RADIO — deep dive scripts

Backpressure narrative: Ticks arrive faster than React can commit. Keep only the latest tick in a ref; schedule one requestAnimationFrame to flush to state/chart. Optionally downsample series to the chart’s pixel width (LTTB). Memory stays bounded with a ring buffer. Per-widget error boundaries keep one bad panel from blanking the page.
A11y narrative for dense data: Every chart needs a tabular alternative or downloadable data. Legends are keyboard toggles, not colour-only swatches. Alert tables are virtualized with row focus and aria-rowcount when known. Live updates use polite status (“alerts updated”) rather than assertive spam.
text
1// Build order for a 45m chat design2// 1. Clarify 1:1 vs group, offline, media, scale3// 2. NFRs: optimistic latency, a11y, INP, CLS on history4// 3. Boxes: composer, list, connection manager, API, presence5// 4. State buckets map6// 5. Deep dives: optimistic + scroll + reconnect (pick 2–3)7// 6. Metrics + flag + risks8// 7. What you'd ship in week 1 vs later

End-to-end chat mock script (45 minutes)

Minute 0–5: clarify 1:1 vs groups, platforms, offline, media, multi-device, expected message rate on hot channels. Minute 5–10: NFRs — optimistic send feel, a11y live regions without spam, INP on composer, CLS on history prepend, delivery at-least-once with idempotent clientId. Minute 10–20: draw client boxes (composer, virtualized list, connection manager, outbox) and server boxes (msg service, presence, media); map four state buckets. Minute 20–35: deep-dive two hard problems — almost always optimistic send and scroll anchoring; add reconnect if time. Minute 35–45: metrics, feature flag, multi-tab socket leadership, week-1 build slice vs later.
Dashboard mock script: clarify live vs historical, tick rate, widget count, alert volume, a11y needs. Pick SSE for one-way ticks unless bidirectional control dominates. Pipeline: network → normalize → downsample → ring buffer → chart with rAF backpressure. Virtualize alert table. Per-widget loading/error. Colourblind-safe series + data table alternative. Close with dropped-tick metrics and a kill switch for live mode.
  1. 01Week-1 chat slice — text-only 1:1, HTTP history + WS fanout, optimistic send, basic virtualize.
  2. 02Week-2 — scroll anchoring on prepend, reconnect catch-up, typing throttle.
  3. 03Week-3 — offline outbox, multi-tab leader, read receipts merge.
  4. 04Later — media, threads, search, E2E encryption if required.
  5. 05Anti-goals — designing Cassandra partitions for 30 minutes in an FE round.
text
1// Catch-up after resume2ws.send(JSON.stringify({ type: 'resume', channelId, lastSeq }))3// server: events where seq > lastSeq, then live4// client: append by seq; Map by id for de-dupe5// if gap detected: HTTP GET history?after=lastSeq as fallback
docsMDN — WebSocketsMDNdocsMDN — Server-sent eventsMDNarticleweb.dev — Virtualize large listsweb.devarticleGreatFrontEnd — Chat ApplicationGreatFrontEndarticleAbly — WebSockets vs SSEAbly

Checkpoint

In chat design, what is the strongest reason to virtualize the message list?

ATo improve SEO of messages.BLong threads create thousands of DOM nodes — virtualization keeps memory and INP sane.CSo WebSockets can reconnect faster.
Sign up free to answer and see why

Checkpoint

User scrolls up to read history; new page of older messages loads. Viewport jumps. What did you miss?

AHTTPS.BScroll anchoring / scrollTop compensation when prepending content.CUsing REST instead of GraphQL.
Sign up free to answer and see why

Checkpoint

Dashboard ticks arrive faster than the chart paints. Best client strategy?

AIncrease React state updates to every tick for accuracy.BCoalesce to animation frames / drop intermediate ticks; downsample to pixel budget.COpen multiple WebSockets to parallelize.
Sign up free to answer and see why

Checkpoint

When is SSE a better default than WebSocket for a live dashboard?

ANever — WebSocket always wins.BPrimarily server→client ticks, simpler reconnection story, no need for client→server on the same socket.CWhen you need binary client uploads on the same channel.
Sign up free to answer and see why

Checkpoint

Optimistic chat send fails. What must the UI guarantee?

ADelete the message silently.BVisible failed state, retry/outbox, and no double-send without idempotency key (clientId).CCrash the tab to avoid inconsistent state.
Sign up free to answer and see why

Could you drive a 45-minute FE system design on chat or a live dashboard without freezing on transport and scroll?

New to itGetting thereConfident

Takeaways

  • Run RADIO on the clock: Requirements → Architecture → Data → Interfaces → Observability.
  • Chat: optimistic send, virtualization, scroll anchoring, reconnect catch-up, multi-device.
  • Dashboard: push path, downsampling, backpressure, per-widget states, contrast.
  • WS vs SSE vs poll is a constraint choice — say the tradeoff.
  • Always budget a11y + perf unprompted.

Next: collab editors, infinite feeds, and autocomplete — harder data/render decisions.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.