Lesson 3 of 4 · 60 min

Design one path before scaling the diagram

Defend a reservation service with a clear invariant, capacity estimate, and failure boundary.

An open-ended design prompt gives you permission to ask questions. It does not require an immediate diagram with every familiar database. First identify the user action, the rule that cannot be broken, and the expected workload. State assumptions when information is unavailable, then show how a changed assumption would change the design.
Separate a business guarantee from an availability target. A booking system may forbid overselling even during a regional failure. That requirement can justify rejecting or delaying writes when the authoritative allocation system is unavailable. Another product may allow a temporary waitlist instead. The candidate's job is to make this trade-off explicit, not to claim that every design provides maximum consistency and availability at no cost.
Build one complete path. Describe how the caller authenticates, how the server validates input, where the invariant is enforced, what commits, and how the user learns the result. Then walk a failure through that path. This exposes missing state more effectively than adding another rectangle. Only after the path works should you discuss partitioning, replicas, or geographic distribution.
Capacity arithmetic should explain a decision. If peak arrival is 200 reservations per second and average write transaction time is 20 milliseconds, average active transactions are roughly four under stable conditions. That does not mean four connections are sufficient for all bursts. It does tell you that a million-way distributed design has no support in the supplied workload.

Worked example

A fictional event sells 5,000 assigned seats. Peak demand is 200 attempts per second. The chosen design uses a single authoritative seat allocation table with unique event/seat ownership and account-scoped idempotency records. The transaction either creates the reservation plus notification intention or rejects the seat conflict. Read replicas can serve event descriptions, but the final allocation decision reaches the writer.
An email worker sends confirmation after commit. Delayed email does not revoke the reservation. If the primary write path is unavailable, the API says reservation status is pending or unavailable according to the operation evidence. It does not accept an untracked claim in a second region. This design prioritizes correct allocation over accepting every request during an outage.

Keep a design decision ledger

A diagram becomes useful when its components connect to explicit requirements. The following original ledger accompanies the assigned-seat design.
RequirementSelected mechanismCost or limit
Never allocate one seat twiceAuthoritative unique event/seat allocationWrites need the authority
Recover a lost replyDurable account/type-scoped operation resultRetention and conflict policy
Email after valid reservationOutbox plus idempotent downstream handlingDelivery can be delayed or repeated
Fast public event browsingSeparately cached public representationFreshness policy required
User can verify purchaseRetrievable reservation by permitted identityAuthorization on reads
The ledger makes it harder to claim that one mechanism solves every requirement. A read replica helps browsing capacity but does not decide final allocation under a no-oversell rule. A queue helps defer email but does not by itself establish one reservation. A cache can reduce read load while creating freshness and access-control questions.

Trace one request through every boundary

code
1caller -> authenticate account A22       -> validate event, seat, expected purchase terms, operation K93       -> authoritative transaction:4            verify replay/conflict5            allocate event E8 seat S76            store reservation B427            store notification intention N68       -> commit9       -> return B42 or recover it after a lost reply10worker -> deliver N6 with stable identity11reader -> retrieve B42 under current permitted access
This trace is more valuable than an unlabeled service diagram because it exposes the atomic boundary. Ask what happens at every arrow if the process stops. Before commit, the operation can be retried without a committed reservation. After commit, the operation result must be recoverable. After notification publication, duplicate delivery must not create a second business effect.

Change the workload rather than the nouns

Suppose demand changes from 200 attempts per second across many events to 20,000 attempts per second for one seat. Partitioning by event ID does not distribute that single hot allocation decision. The system can reject excess attempts early, queue an explicit fair admission process, or use another policy that preserves the seat invariant. The choice should state what the user sees and whether arrival order or a lottery is part of the product contract.
Now suppose the main pressure is 50,000 public page reads per second while reservations remain at 200 attempts. A public cache can address the read workload independently. The final write still reaches authority. This change illustrates why workload shape matters more than a generic claim that the product is high scale.
A geographic requirement can change the availability discussion. If the product must accept reservations during a partition in two regions, it needs a valid allocation strategy, such as preallocating disjoint seat ownership or another protocol with a proven invariant. Allowing both regions to allocate the same seat from stale replicas violates the original rule. Do not claim that adding replication automatically resolves this trade-off.

A cross-row capacity variation

For unassigned capacity, use a counter or other serialized allocation decision rather than seat uniqueness. If the model spreads allocations across rows and sums them before approval, the guard/isolation contract must be explicit. At PostgreSQL Read Committed, acquire the same capacity guard before reading the sum and require every writer to follow it. A pre-existing Repeatable Read snapshot does not become fresh merely because the transaction later waits for that guard.
The candidate does not need to choose this design for every database. The important skill is to connect the database's documented semantics to the exact schedule. Serializable transactions can be another valid choice if the complete transaction is retried after a serialization failure.

Misconceptions to correct

The first misconception is that a larger diagram demonstrates more senior judgment. Every new boundary needs a reason, failure behavior, and operator. Unnecessary components can make a small design less reliable and harder to explain.
The second misconception is that scaling reads solves contention on a single write invariant. Replicas can serve eligible reads, but one scarce resource still needs a correct allocation decision. Identify whether the bottleneck is broad traffic or concentrated contention.

Extend the exercise

Add a requirement that confirmation email may be delayed for ten minutes but a reservation result must be retrievable immediately after commit. Explain which path can degrade without breaking allocation. The model answer allows asynchronous notification delay while preserving the authoritative reservation lookup. Award one point for separating effects, one for a truthful user status, and one for monitoring the delayed notification backlog.

Exercise and solution

The interviewer changes the requirement to unassigned seating with capacity 5,000. Explain the change. A unique seat key no longer models capacity. Use a conditional counter allocation or another serialized capacity decision, and keep reservation identity in the same transaction. Award one point for recognizing the changed invariant, one for a valid capacity control, and one for preserving retry safety.

Interview probe and wrap-up

What would make you partition the allocation store? A strong answer refers to measured contention, independent events, geographic requirements, or operational limits. Follow up with one extremely popular event that cannot be split by event ID. A weak answer partitions because distributed systems are expected in interviews. End with the guarantee, its cost, and the observation that would justify the next level of complexity.

Sources

docsPostHog technical-screen guidanceposthog.comdocsPostgreSQL transaction isolationpostgresql.orgdocsAWS transactional outbox patterndocs.aws.amazon.com

Checkpoint

What does a public read cache directly help in this design?

AHigh-volume eligible event browsing under a freshness policy.BAtomic notification delivery.COperation-key retention after expiry.DFinal seat allocation correctness.
Sign up free to answer and see why

Checkpoint

20,000 attempts target one seat. Why may event-ID partitioning fail to distribute the decision?

APartitioning increases capacity equally for every popularity distribution.BMore partitions let contenders for the same seat decide independently.CA read replica removes contention at the allocation writer.DThe hot event still maps to one allocation authority.
Sign up free to answer and see why

Checkpoint

Two isolated regions both allocate the same seat from stale copies. Which claim fails?

ARead availability can improve.BRegions can operate independently.CThe original no-oversell invariant is preserved automatically.DPublic descriptions can be cached.
Sign up free to answer and see why

Checkpoint

Email is delayed after reservation commit. What can remain correct?

AThe retrievable authoritative reservation, with notification pending.BA response that reports the reservation as failed until email arrives.CA response that reports email delivery from the reservation commit alone.DA retry that allocates a second seat because the first email is delayed.
Sign up free to answer and see why

Checkpoint

A guard is added after a Repeatable Read snapshot was established. What should the candidate avoid claiming?

AThat every writer must follow a common protocol.BThat the lock necessarily refreshes the earlier snapshot.CThat a shared decision needs protection.DThat retries may be necessary.
Sign up free to answer and see why

Can you connect each mechanism to one requirement, then revise the design for a hot seat, read-heavy traffic, or a geographic partition without silently weakening allocation correctness? State the relevant identifiers, failure boundary, and evidence in your own words before selecting your confidence.

Not yetGetting thereConfident

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.