Lesson 3 of 4 · 35 min

A committed order needs a durable next step

Close the database-to-message gap and explain duplicate delivery.

An order service often needs two effects: save the order and tell another service to act. These effects usually live in different systems. Sending the message first can announce an order that never commits. Committing first and sending next can lose the announcement when the process crashes. Changing the order of these two lines does not remove the gap.
The transactional outbox makes the intention to send part of the same database transaction as the order. A separate dispatcher reads committed intentions and publishes them. The transaction creates either both records or neither. This provides a durable connection between business state and work that remains to be delivered. It does not make the network publish and the dispatcher acknowledgement one atomic operation.
The dispatcher can publish a message and crash before marking it sent. Its next attempt sends the same message again. The consumer therefore needs a stable event identifier and an idempotent effect. When both its deduplication record and local business update use one transaction, a duplicate can become a harmless replay. If the effect is external, such as sending a bank transfer, local deduplication alone is insufficient. The external provider also needs an appropriate identity or reconciliation contract.
Ordering requires a separate decision. An event identifier tells you whether two deliveries refer to one event. A sequence number tells you where distinct events belong within an aggregate. Do not assume timestamps provide total ordering. Clocks can disagree, and two events can share a timestamp. Many workflows can process independent orders concurrently while preserving order only inside one order.

Worked example

Order O8 moves from draft to confirmed at version 3. The same transaction inserts event E31 with aggregate O8, version 3, and type order-confirmed. The dispatcher publishes E31, then crashes. The warehouse receives E31 and transactionally inserts processed-event E31 plus allocation A9. On retry it receives E31 again. The uniqueness check finds E31 and the warehouse performs no second allocation. There are two deliveries and one business effect.
Now imagine version 4 cancellation arrives before version 3 confirmation. A consumer cannot simply discard whichever message arrives second. It must buffer a gap, fetch current authoritative state, or use a state transition rule that safely handles the order. State which policy the workflow requires.

Inspect the two atomic boundaries

The following records describe an original teaching design. They are not a claim about a particular queue's configuration. The producer stores the order and outbox event in one local transaction. The consumer stores its processed-event marker and allocation in another local transaction.
code
1producer order:2  orderId=O8, account=A2, status=confirmed, version=33producer outbox:4  eventId=E31, aggregateId=O8, aggregateVersion=35  type=order-confirmed, deliveryState=pending6consumer processed event:7  consumer=warehouse, eventId=E318consumer allocation:9  allocationId=A9, orderId=O8, sourceEvent=E31
The producer transaction cannot directly guarantee the consumer transaction. The queue and network sit between them. The design instead makes unfinished delivery recoverable and repeated delivery safe at the consumer's local effect. Notice that event identity and order identity both appear. They support different deduplication needs.
If two distinct events describe the same order confirmation, uniqueness only on eventId can still permit two allocations. The allocation model may also need one active allocation per order. The exact rule depends on whether an order can legitimately produce multiple allocation events. A consumer must enforce the business effect rule, not assume the publisher's event format makes the rule unnecessary.

Walk the crash matrix

Crash locationProducer stateConsumer stateRecovery
Before producer commitNo order/eventNoneRetry original operation
After commit, before publishOrder plus pending E31NoneDispatcher retries E31
After publish, before sent markOrder plus pending E31Maybe processedReplay E31 safely
After consumer commit, before ackProducer may mark sentAllocation plus E31 markerDuplicate becomes no-op
This matrix separates committed state from acknowledgements. A consumer acknowledgement is a message to the queue, not the business effect itself. A sent mark is local dispatcher state, not proof that the downstream user has received a final benefit. Each layer should name what its completion flag actually means.
Now consider a dispatcher that marks an event sent before publishing to avoid duplicates. A crash between those steps loses the message permanently unless another mechanism detects and repairs the gap. That change replaces duplicate risk with loss risk. Since duplicate delivery can be handled through stable identity, publishing before marking sent is often the recoverable sequence for this pattern.

Ordering and schema changes

An aggregate version lets a consumer detect a missing predecessor. If version 4 arrives before version 3, choose a documented policy. A projection that only needs current status can fetch the authoritative order and replace its view. A financial sequence that needs every transition may require buffering and replay. Neither policy should be inferred from event arrival order.
Schema evolution adds another failure path. A consumer may receive an event version it does not understand. Keep the event identity and payload or safe payload reference available for diagnosis, classify the failure, and avoid an endless tight retry loop. A repair can deploy a compatible consumer and replay the same event. Generating a new event ID during replay may defeat deduplication and conceal the relationship to the original failure.
An outbox table also needs operational care. Index pending work, bound polling batches, monitor oldest pending age, and define retention after safe delivery. These are design requirements for the teaching example, not instructions to deploy a particular database configuration. Deleting pending records to reduce table size is data loss, while retaining every delivered payload forever can create unnecessary storage and privacy burden.

Misconceptions to correct

The first misconception is that a queue's delivery claim covers every external effect. Even when a broker offers stronger delivery features, a consumer can crash after sending an email or calling a provider but before acknowledging. The external effect still needs its own contract.
The second misconception is that an outbox is the same as event sourcing. An outbox stores delivery intentions tied to current business changes. Event sourcing uses an event history as the primary model of business state. A system can use an outbox without reconstructing all state from events.

Extend the exercise

A consumer receives E31 twice and E32 once, all for O8 confirmation. Specify which uniqueness checks stop duplicate allocations and which records remain for audit. A strong answer tracks both event deliveries while protecting the one-allocation-per-order rule. Award two points for the distinct identities and one for retaining a replayable failure history.

Exercise and solution

The order transaction rolls back after the outbox insert. Should a dispatcher publish the row? No, the insert rolls back with the order and is invisible as committed work. Next, the consumer crashes after allocation but before acknowledgement. Its transaction has committed, so a duplicate delivery finds the deduplication record. Award one point for atomic order/outbox state, one for duplicate handling, and one for distinguishing acknowledgement from business completion.

Interview probe and wrap-up

Does an outbox guarantee exactly-once execution everywhere? A strong answer limits the claim to atomic local state plus eventual delivery attempts, with idempotency at later effects. Ask what happens during permanent poison events. A weak answer says the queue removes all duplicates. Keep each guarantee attached to the boundary where it actually holds.

Sources

docsAWS transactional outbox patterndocs.aws.amazon.comdocsStripe webhook delivery and verificationdocs.stripe.com

Checkpoint

What must commit together at the producer?

AThe order and its outbox intention.BThe browser response and queue ack.CEvery downstream service transaction.DThe queue delivery and email receipt.
Sign up free to answer and see why

Checkpoint

A dispatcher marks sent before publication and crashes. Main risk?

AAutomatic consumer rollback.BA duplicate allocation already committed by the consumer.COnly a harmless duplicate.DA lost delivery intention unless another recovery mechanism exists.
Sign up free to answer and see why

Checkpoint

E31 and E32 both describe one order confirmation. Event-ID deduplication alone may miss what?

ATwo transport attempts for E31.BAn uncommitted outbox insert.CA repeated business allocation for the same order.DA delayed queue acknowledgement for the same delivery.
Sign up free to answer and see why

Checkpoint

Version 4 arrives before version 3. Which policy is defensible?

AUse a documented gap policy such as replay or authoritative-state refresh.BAlways process arrival order as business order.CDiscard version 3 without knowing the consumer's needs.DCompare timestamps only and assume total order.
Sign up free to answer and see why

Checkpoint

Why is an outbox not automatically event sourcing?

AAn outbox's sent flag is itself the complete authoritative business history.BThe outbox can store delivery intentions while current records remain authoritative.CAn outbox always requires replaying its events to reconstruct current business state.DEvent sourcing is defined only by using a message broker.
Sign up free to answer and see why

Can you use the crash matrix to identify which local transaction committed, then explain why a duplicate event and a duplicate business effect need separate reasoning? State the relevant identifiers, failure boundary, and evidence in your own words before selecting your confidence.

Not yetGetting thereConfident

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.