Communication
Design Reminder Alert
01
Requirements
Requirements
- Create a one-time reminder for a specific datetime + timezone
- Create a recurring reminder (daily, weekly, custom RRULE)
- Deliver via push notification, SMS, or email — user configures channel
- Snooze (reschedule N minutes/hours from now)
- Cancel or edit a pending reminder
- Support all IANA timezones including DST transitions
- Delivery within ±30 seconds of the scheduled time
- At-least-once delivery — never silently drop a reminder
- Client-side dedup so user sees it once even if delivered twice
- Scale to ~100M reminders/day = ~1200/sec avg, ~5K/sec peak
- 99.99% availability on the scheduling path
- Recurring reminders must handle DST, leap years, "last day of month"
02
Scale Estimation
Scale Estimation
03
API Design
API Design
Create reminder. Body: {user_id, title, body, fire_at: "2026-04-17T09:00", timezone: "America/New_York", channel: "push", recurrence?: "RRULE:FREQ=DAILY"}. Server stores fire_at_utc computed from local time + IANA zone. Returns {reminder_id}.
List user's pending reminders. Paginated by fire_at_utc. Includes next fire time for recurring.
Snooze: reschedule to now + N minutes. Creates a new one-time instance; original recurrence continues separately.
Cancel pending reminder. For recurring: cancels the series; individual instance cancel via PATCH.
04
Architecture
Architecture
Two flows: scheduling (user creates → store → index by fire_at_utc) and firing (scanner reads due reminders → dispatch to delivery channels). A recurrence expander materializes the next N instances of recurring reminders.
05
Deep Dive — Time Buckets + Timezone + At-Least-Once
Deep Dive — Time Buckets + Timezone + At-Least-Once
Time-bucketed scanning. Reminders stored in Cassandra with partition key = fire_minute_utc (truncated to the minute). E.g., all reminders due at 2026-04-17T13:00 UTC are in partition 202604171300. Scanner wakes every 60 s, reads the current-minute partition, dispatches each reminder.
Why this works: partition key gives O(1) lookup of "everything due now." No full-table scan. Scanner is a single leader process (elected via distributed lock) per shard of the time space. Multiple scanners shard by minute-range to parallelize.
Timezone materialization. User says "9am in America/New_York." Server must compute the UTC equivalent at creation time. But: DST changes mean "9am" maps to different UTC offsets on different dates. For one-time reminders, compute once at creation. For recurring: compute the NEXT instance's UTC at recurrence-expand time, re-compute after each fire.
Critical rule: store the IANA zone name (America/New_York), NOT the UTC offset (-05:00). UTC offsets change with DST. If you stored offset, a daily 9am reminder set in January (-05:00) would fire at 10am local time in March when clocks spring forward (-04:00).
sequenceDiagram
participant SC as Scanner
participant DB as Reminder DB
participant DD as Dedup (Redis)
participant D as Dispatcher
participant CH as Push / SMS / Email
participant RE as Recurrence expander
SC->>DB: read partition 202604171300
DB-->>SC: [reminder_1, reminder_2, ...]
loop each reminder
SC->>DD: SETNX reminder_id (dedup)
alt new (not seen)
DD-->>SC: 1 (acquired)
SC->>D: dispatch reminder
D->>CH: deliver via configured channel
CH-->>D: ack / fail
alt recurring
D->>RE: compute next instance
RE->>DB: insert next fire_minute_utc partition
end
D->>DB: mark delivered
else already processed
DD-->>SC: 0 (skip)
end
end
At-least-once delivery. Scanner reads due reminders → dispatches → marks delivered. If scanner crashes after dispatch but before marking: next scanner run sees the same reminder, re-dispatches. Hence at-least-once. The dedup store (Redis SETNX with TTL) prevents most duplicates; client-side dedup (by reminder_id) catches the rest.
Recurring expansion. After a recurring reminder fires, the recurrence expander computes the next instance from the RRULE + IANA timezone, converts to UTC, and inserts into the appropriate time-bucket partition. Horizon: expand only 1–2 instances ahead (not "every Monday forever").
"Store reminders in Cassandra partitioned by fire_minute_utc. Scanner (leader-elected per shard) wakes every 60 s, reads the current-minute partition, dedup-checks via Redis SETNX, and dispatches to the configured channel (push/SMS/email). At-least-once by design — scanner re-processes unacked reminders. Client dedup by reminder_id. Timezones stored as IANA zone names, not offsets; UTC conversion happens at creation + recurrence expansion. Recurring reminders expand one instance ahead after each fire."
06
Anti-patterns
Anti-patterns
DST changes the offset twice a year. A daily 9am reminder created in winter fires at 10am local time in summer.
At 100M reminders, scanning the whole table every 60s is ~1.7M rows/sec just for the scan. Unscalable.
Infinite storage. And if the user edits the recurrence, you have to delete + recreate all future instances.
07
Tradeoffs & Design Choices
Tradeoffs & Design Choices
- Polling (scanner every N sec) vs timer-wheel / DelayQueue. Polling is simpler and works well for minute-level precision. Timer wheels (Kafka-style) are better for sub-second precision but harder to distribute. For reminders (30 s tolerance): minute-bucket polling wins on simplicity.
- At-least-once vs exactly-once delivery. Exactly-once across distributed systems is infeasible without 2PC (slow) or app-level dedup. At-least-once + dedup is the pragmatic answer. Client-side dedup is cheap (check reminder_id before displaying).
- Cassandra vs Postgres for time buckets. Cassandra: partition by fire_minute_utc gives free sharding + fast partition reads. Postgres: range query on fire_at index, simpler ops. At 100M/day, Cassandra is the better fit. At 1M/day, Postgres is fine.
- Single scanner leader vs sharded scanners. Single: simpler, no coordination needed, but SPOF. Sharded: each scanner owns a range of minutes (scanner_0 handles even minutes, scanner_1 handles odd, etc.). Leader election per shard via Redis lock.
- Push vs pull for delivery confirmation. Push to APNs/FCM is fire-and-forget (no delivery guarantee from the platform). SMS has delivery receipts. Email has no reliable read-receipt. Accept: delivery = "we sent it"; display confirmation is the client's problem.
08
Failure Modes
Failure Modes
09
Evolution
Evolution
MVP — cron job + single Postgres
Cron scans WHERE fire_at <= NOW() AND status = 'pending' every minute. Sends email. Works to ~10K reminders/day.
Time-bucketed Cassandra + scanner leader
Partition by fire_minute_utc. Scanner leader per shard. Multi-channel delivery (push + SMS). ~1M/day.
RRULE engine + IANA timezone handling
Proper recurring support. Zone-aware UTC conversion. DST-safe. ~10M/day.
Channel escalation + delivery tracking
Push fails → escalate to SMS after 5 min. Delivery log for audit. At-least-once with Redis dedup. ~100M/day.
Smart reminders + ML-suggested timing
ML predicts when user is most likely to act on the reminder. "You usually check email at 8:15am; should we remind at 8:10?" Context-aware delivery.
Watch and read
Try next
Free to read · better with Enzo
Whiteboard this with Enzo
Enzo runs it as a live system design round on the whiteboard and grades your trade-offs.