Infrastructure
Design Unique ID Generator
01
Requirements
Requirements
- Generate globally unique 64-bit IDs from any machine, no central bottleneck
- IDs must be k-sortable (roughly ordered by creation time)
- Support batch generation: request N IDs in a single call
- Parse an ID back into its components: timestamp, machine_id, sequence
- Encode IDs as compact base62 strings (~11 chars) for URL use
- Sub-millisecond latency per ID generation (local operation)
- Throughput: 4,096 IDs/ms per machine, 1M+ globally
- No single point of failure for ID generation itself
- IDs must not leak sensitive info (exact server count, internal topology)
- Monotonically increasing within a single machine
- Survive NTP clock adjustments without producing duplicates
02
Scale Estimation
Scale Estimation
Each machine can produce 4,096 IDs per millisecond (12-bit sequence). With 1,024 machines (10-bit machine_id), the system yields 4,096 x 1,024 x 1,000 = ~4.2 billion IDs/sec theoretical max. In practice, 1M+/sec sustained is the target.
A 41-bit timestamp in milliseconds covers 2^41 ms = ~69.7 years from a custom epoch. IDs are 64 bits = 8 bytes; base62-encoded they become ~11 characters.
03
API Design
API Design
Generate N unique IDs (default 1, max 1000). Returns array of 64-bit IDs as both integer and base62 string. Each ID encodes timestamp + machine_id + sequence.
Decode a 64-bit ID (integer or base62) into its components: { timestamp, machine_id, sequence, created_at_utc }. Useful for debugging and audit.
Returns generator status: current machine_id, last timestamp, sequence position, clock health (detects backward jumps).
04
High-Level Architecture
High-Level Architecture
The ID service is stateless per-request — each instance holds only its machine_id and a local clock + sequence counter. No database, no network call on the hot path.
- ID Service instances: one per application host (or sidecar). Generates IDs locally with zero network overhead.
- Zookeeper / etcd: assigns unique machine_ids on startup. Only needed at boot, not on the hot path.
- Alternative — Ticket Servers: dual-MySQL with odd/even auto-increment (Flickr's design). Centralized but simpler; acceptable at lower scale.
05
Deep Dive — Snowflake Bit Layout & Alternatives
Deep Dive — Snowflake Bit Layout & Alternatives
How do you pack enough information into 64 bits to guarantee global uniqueness, preserve time-ordering, and avoid any coordination between machines on the hot path?
Snowflake (Twitter, 2010). The canonical design: 1 sign bit + 41-bit timestamp (ms since custom epoch, ~69 years) + 10-bit machine_id (1,024 machines) + 12-bit sequence (4,096 IDs per ms per machine). Total: 64 bits. IDs are k-sortable because the timestamp occupies the most-significant bits. Each machine generates IDs independently; the machine_id guarantees no collisions across hosts.
ULID (Universally Unique Lexicographically Sortable Identifier). 128 bits: 48-bit timestamp (ms, ~8,919 years) + 80-bit cryptographically random tail. No machine registration needed — randomness provides uniqueness. Tradeoff: slight collision probability (birthday problem at ~2^40 IDs/ms, negligible in practice). Encodes to 26 chars Crockford base32.
UUID v7 (RFC 9562). 128-bit, timestamp-first layout for B-tree index friendliness. Replaces UUID v4's fully-random approach with time-ordered prefix. Wider than Snowflake (128 vs 64 bits) but standardized and needs no coordinator.
Ticket Servers (Flickr, 2010). Two MySQL instances: one auto-increments by 2 starting at 1 (odd), the other starting at 2 (even). App round-robins between them. Simple, no clock dependency, but centralized — each MySQL is a SPOF risk. Flickr ran this in production for years.
Clock skew handling. Snowflake relies on the local clock. If NTP corrects the clock backward, the generator would produce IDs with a past timestamp — risking duplicates. Defense: if current_ms < last_ms, either refuse to generate (throw error) or spin-wait until the clock catches up. Critical invariant: never generate an ID with a timestamp older than the last generated ID.
sequenceDiagram
participant App as Application
participant Gen as ID Generator
participant Clk as System Clock
App->>Gen: generate_id()
Gen->>Clk: current_time_ms()
Clk-->>Gen: timestamp
alt same ms as last call
Gen->>Gen: sequence++
alt sequence > 4095
Gen->>Gen: spin-wait for next ms
Gen->>Clk: current_time_ms()
Clk-->>Gen: new timestamp
Gen->>Gen: sequence = 0
end
else new ms
Gen->>Gen: sequence = 0
end
Gen->>Gen: compose: sign|timestamp|machine|seq
Gen-->>App: 64-bit ID
App->>App: base62_encode(id) for URL use
06
Anti-patterns
Anti-patterns
Single DB becomes a write bottleneck and SPOF. IDs are sequential and predictable — attackers can enumerate all records. Sharding breaks the sequence.
Fully random UUIDs fragment B-tree indexes — every insert goes to a random leaf page, causing excessive page splits and write amplification. 128 bits is also twice as wide as needed, wasting index space.
Two machines generating an ID in the same millisecond produce identical values. Even on one machine, burst traffic within 1 ms causes collisions.
07
Key Design Decisions & Tradeoffs
Key Design Decisions & Tradeoffs
Timestamp + machine_id + sequence in 64 bits
K-sortable, compact, 4,096 IDs/ms/machine. Requires a machine_id registry (Zookeeper/etcd) at boot time. Clock-dependent — NTP backward jump must be handled. The gold standard for high-throughput systems (Twitter, Discord, Instagram).
Timestamp + random tail, no coordinator
No machine registration needed — randomness provides uniqueness. Wider (128 bits, 26 chars). Slight theoretical collision risk under extreme load within same ms. Better for systems that cannot run Zookeeper.
Each machine generates IDs independently
Sub-microsecond latency; no SPOF on the hot path. Machine_id assigned once at startup. The Snowflake/ULID approach. All production systems use this.
Dual-MySQL auto-increment (Flickr-style)
Simpler to reason about; strictly ordered. But every ID requires a network round-trip to the DB. Two servers with odd/even offset provide redundancy but not horizontal scale. Works at moderate throughput only.
Start timestamp from a recent date, not Unix epoch
41-bit ms from Jan 1, 2020 lasts until ~2089. Starting from Unix epoch (1970) would waste 50 years of bits, reducing usable lifespan to ~19 years from now.
Standard milliseconds since 1970
Universal, no custom epoch to document. But wastes precious timestamp bits on the past. Only viable with wider formats (ULID's 48 bits = 8,919 years).
08
What Can Go Wrong
What Can Go Wrong
09
Interview Tips
Interview Tips
- Start with requirements, not solutions. Ask: how many IDs/sec? Must they be sortable? How compact? 64-bit or 128-bit acceptable? This shapes the entire design.
- Draw the bit layout. Literally sketch
1 + 41 + 10 + 12 = 64on the whiteboard. Interviewers love seeing the tradeoff between timestamp range, machine count, and per-ms throughput. - Address clock skew proactively. Mention NTP backward jump before being asked. The defense (refuse or wait) shows you understand the real-world failure mode that trips up Snowflake.
- Know the alternatives. Snowflake vs ULID vs UUID v7 vs Ticket Server. Each has a sweet spot. Naming all four and their tradeoffs demonstrates breadth.
- Mention base62 encoding. 64-bit integer to ~11-char URL-safe string. Shows you think about the consumer of IDs, not just the generator.
10
Evolution
Evolution
Auto-increment database column
Single MySQL/Postgres with BIGINT AUTO_INCREMENT. Simple, strictly ordered. Works to ~10K writes/sec on one machine. Cannot scale horizontally without splitting ranges.
UUID v4 (random)
No coordination, no central DB. But 128 bits (too wide), not sortable, fragments B-tree indexes. Good enough for low-write systems that don't need ordering.
Ticket servers (Flickr, 2010)
Dual-MySQL with odd/even auto-increment offsets. Round-robin between them for redundancy. Centralized but practical. Flickr used this for all photo IDs in production.
Snowflake (Twitter, 2010)
The breakthrough: 64-bit, k-sortable, fully distributed. Each machine generates independently. Became the industry standard. Used by Twitter, Discord, Instagram (with variations).
ULID / UUID v7 (modern, no coordinator)
Timestamp-first + random tail. No Zookeeper needed. UUID v7 standardized in RFC 9562 (2024). Best choice for new systems that want sortability without infrastructure overhead.
Watch and read
References & Videos
Try next
Free to read · better with Enzo
Whiteboard this with Enzo
Enzo runs it as a live system design round on the whiteboard and grades your trade-offs.