Explain delivery guarantees at the sink transaction boundary.
Mechanism and reasoning
A broker's transaction guarantee has a scope. It may coordinate consumed offsets and produced broker records, yet an external database or HTTP call remains outside that transaction unless the design explicitly includes it. Saying 'exactly once' without drawing the boundary hides the most important failure case.
Consider a consumer that writes a result and then commits its offset. If it crashes between those steps, the broker can redeliver the event and the write can happen again. Reversing the order avoids that duplicate window but creates a loss window: the offset can advance before the result is durable.
An idempotent sink makes repeated processing converge to the same business result. One approach stores the event ID and business effect in the same database transaction under a uniqueness constraint. A retry observes the ID and skips the repeated effect. The transaction must be atomic across both records. A separate 'processed' flag written earlier does not provide the same guarantee.
Not every effect is naturally idempotent. Setting an order's status to a versioned value is easier than incrementing a balance or sending a notification. External calls need their own idempotency contract or a durable outbox and reconciliation process. An outbox itself still needs duplicate-safe delivery.
In interviews, walk through crash points. Show state before the transaction, after commit and before the broker offset update. Then explain how restart behaves. This is stronger than repeating a delivery label. It also clarifies retention: deduplication records cannot expire earlier than the maximum redelivery horizon without reopening the duplicate risk.
Walk every crash boundary with a ledger
Use a balance update to make the failure visible. The business operation is a credit identified by a stable operation ID. The sink must either commit the marker and effect together or commit neither. A uniqueness constraint is the arbiter when multiple consumers race on the same ID; a prior read alone is not sufficient because both consumers can observe absence.
sql
1-- PostgreSQL-style teaching sketch; marker and credit share one transaction.2BEGIN;3WITH accepted AS (4 INSERT INTO processed_event(event_id)5 VALUES ('credit-42')6 ON CONFLICT (event_id) DO NOTHING7 RETURNING event_id8)9UPDATE account10SET balance = balance + 1011WHERE account_id = 'a7'12 AND EXISTS (SELECT 1 FROM accepted);13COMMIT;
This sketch assumes the account exists, the amount and account are validated, and the event ID is bound to the intended immutable operation. A production implementation must handle missing accounts, conflicting payloads under the same ID, retries after uncertain commit responses, permissions and observability. If the marker commits but the account update affects no row unexpectedly, the operation could be marked complete without applying the intended effect. Validate that outcome within the transaction or use a schema and procedure that enforce the contract.
Crash point
Sink state
Broker may redeliver?
Recovery
Before sink transaction
No effect, no marker
Yes
Apply transaction
During uncommitted transaction
Rolled back under stated DB semantics
Yes
Apply transaction
After sink commit, before offset commit
Effect and marker durable
Yes
Detect duplicate, no second credit
After both commits
Effect and marker durable
Normally no for this position
Replays still use same identity
This table does not claim that each line executes once. The handler may run multiple times. The property is that repeated processing converges to one accepted business effect within the sink boundary and retention horizon. That distinction makes retry behavior easier to explain and test.
An external API creates an uncertain-response case. The consumer sends a payment request; the remote service applies it; the response is lost. Retrying with a new operation key can create a second payment. Retrying with the same supported idempotency key may return the first result, subject to that provider's retention and parameter rules. If the remote API offers no idempotent contract, a durable local outbox alone cannot prove one external effect. Reconciliation or a compensating business process may be needed.
An outbox solves a different atomicity problem: the database can commit the business state and an intent to publish together. A relay later sends the intent. If it crashes after sending but before marking the outbox item delivered, it can send again. Downstream deduplication is still required when repeated delivery is possible. Do not describe the outbox as magic end-to-end exactly-once delivery.
Retention determines the time boundary of the guarantee. If the source can replay thirty days but markers expire after seven, an old credit can become new again to the sink. Keeping a permanent business operation ledger or making the final state update naturally version-idempotent can change that design. The choice depends on the business effect and legal retention constraints. State the horizon explicitly in the contract and test the oldest supported replay.
For verification, run concurrent duplicate deliveries and inject failures around transaction commit and offset commit. Inspect both the marker table and the balance. A final balance alone can hide a missing operation offset by an extra duplicate credit elsewhere. Compare each operation ID with its expected effect and confirm that invalid operations do not leave a success marker.
The interview answer should finish by naming the boundary: “Within this database transaction and while operation identities are retained, repeated delivery produces one credit. The broker can still redeliver, and external calls require their own contract.” This is a precise guarantee that another engineer can challenge with a crash scenario.
Worked example
Teaching transaction for event e42 credits ten units. Initially the balance is 100 and e42 is absent. In one database transaction, insert e42 into a unique processed-events table and add ten to the balance. Commit produces balance 110 and the event marker. The consumer crashes before offset commit. On redelivery, the unique event insert finds e42 already present, so the credit is not repeated. Balance remains 110. This depends on both writes sharing one transaction and on the marker remaining available.
Exercise
A consumer first increments the balance, then writes its deduplication marker in a separate transaction. It crashes between them. Starting balance is fifty and the credit is five. What can replay produce, and how do you fix the boundary?
Model solution and rubric
The first increment produces fifty-five. With no durable marker, replay can increment again to sixty. Put the marker and increment in the same transaction, with a unique event key and correct handling of the duplicate-key path. Alternatively use a sink operation whose business effect is itself idempotent. Test each crash point. Broker offset timing alone cannot make two independent sink transactions atomic.
Score out of four: one point for the correct result, one for showing the intermediate reasoning, one for identifying the stated failure case, and one for a verification that could disprove the answer. Do not award the reasoning point for a tool name alone.
Failure modes and misconceptions
“An outbox prevents duplicate delivery.” It preserves durable publication intent, but a relay can resend after an uncertain acknowledgement. The receiver still needs duplicate-safe behavior.
“Read-before-write deduplication is enough.” Concurrent consumers can both read absence. Enforce uniqueness and the effect in one atomic boundary.
Interview probe
Evidence class: recommended. Original practice.
Does Kafka exactly-once processing guarantee that an external payment API is called once?
Strong answer: No. The external API must participate through its own idempotency or transaction contract. I would use a stable business operation key, durable intent, retry handling and reconciliation for uncertain responses.
Follow-up: How long must deduplication state remain if replay can cover thirty days?
Weak answer indicators: Treating broker transactions as universal; committing offsets before durable effects; expiring markers without considering replay.
Sources
Technical references: Kafka 4.1 design; Airflow best practices. Sources support the documented mechanisms. The numbers, decisions, rubrics and interview prompts in this lesson are original teaching examples, not measurements or employer question claims.
A sink credit commits before offset commit, then the consumer crashes. What protects the credit on redelivery?
AA longer client timeoutBA marker written later in another transactionCA unique marker committed atomically with the effectDCommitting the offset first on every retry
An outbox relay sends successfully but crashes before recording delivery. What may happen?
AThe database rolls back the earlier business transactionBThe receiver cannot see the first messageCNo retry is needed under any protocolDThe relay can send again, requiring duplicate-safe consumption
Two consumers read no marker for the same credit, then each increments the balance in a separate transaction. Which repair closes the race?
AUse a shared durable uniqueness constraint with the marker and credit in one transactionBRead the marker twice before each separate incrementCCommit the marker first, then increment without a recovery state machineDCommit the broker offset before checking the sink
Replay horizon is 30 days; deduplication markers last 7. What limit is exposed?
AAll current-day effects will be lostBOld non-idempotent effects can repeat after marker expiryCThe broker automatically extends marker retentionDExactly-once transport prevents the mismatch
An external API applies a request but its reply is lost. What is the safest retry basis when supported?
AA new operation ID for each retryBThe client's latest timestamp onlyCThe same parameter-bound idempotency key and reconciliation policyDNo durable record of the attempt
Explain how you would explain delivery guarantees at the sink transaction boundary without reading the solution. State one assumption that could change your answer, and one observation that would make you revise it.
Not yetGetting thereConfident
Wrap-up
Name the transaction boundary and simulate a crash on each side. Guarantee the business result, not a slogan.