Lesson 4 of 4 · 25 min

Make replay a controlled operation

Plan a deterministic replay with bounded impact.

Mechanism and reasoning

Replay is a data repair tool, but it can also repeat damage. Before replaying, identify the input range, code version, schema version, destination and external side effects. A consumer that rebuilds a table safely may still resend notifications or duplicate API calls. Separate analytical recomputation from outward effects.
Keep raw input immutable where practical and allowed by retention policy. A replay that reads changing source data is not reproducing the same computation. Record source partitions or snapshots and a run identity. If privacy deletion changes retained input, document that replay reflects the authorized remaining data rather than pretending it is an exact historical reconstruction.
Choose an output strategy. Rebuild into an isolated destination, validate it, then switch a pointer or perform an atomic replacement where the storage system supports it. In-place updates can work with a careful idempotent design, but readers may observe mixed old and new results. Define whether such partial visibility is acceptable.
Validation needs more than total row count. Compare key uniqueness, sums by partition, null rates, expected corrections and a sample of business entities. Equal totals can hide a wrong join or swapped categories. Record expected differences before the run so validation does not become post-hoc explanation.
Bound resource use. A replay can starve the live pipeline or create a broker retention race. Limit concurrency, use a separate consumer identity and monitor live freshness. Pause if a stated threshold is crossed. The interview answer should explain how to stop, resume and prove completion, not only how to launch the job.

Treat a backfill as a versioned data release

A replay manifest should be small enough to inspect before the job starts. It records the input snapshot or positions, transformation revision, parameters, output generation and validation expectations. Use an explicit logical run date when the transformation depends on time. Calling the current clock inside historical logic can make identical input produce different answers tomorrow.
code
1Illustrative replay manifest R172input: immutable source partitions for 2026-09-01 through2026-09-033transform: parser revision P84logical_as_of: 2026-09-04T00:00:00Z5destination: orders_rebuild_R176external_effects: disabled for analytical repair7expected_delta: +2 keys on day 2, −1 invalid key on day 38live_freshness_stop: lag above five minutes for two checks9promotion: validated generation pointer change
The manifest is not proof that the inputs remain accessible. Check retention before starting. If source logs have already expired, the repair may require an older snapshot plus subsequent changes, another authoritative source, or an explicit statement that exact reconstruction is impossible. Increasing retention after the data is gone does not recover it. Document missing ranges instead of quietly treating them as empty.
Build isolation must include side effects. A separate output table prevents readers from seeing partial analytical results, but a transform that sends notifications can still act externally. Disable or redirect those effects under the approved repair contract, or use the same stable operation identities and deduplication horizon as normal processing. A replay-specific consumer group does not by itself make business effects safe.
Validation dimensionExample checkWhy total rows are insufficient
IdentityUnique order keyDuplicates can offset omissions
PartitionCounts and key differences per dayRows can move to wrong dates
AmountSum by currency and dayEqual row counts can hide wrong values
SemanticsThree expected changed keysCorrect totals can arise accidentally
CompletenessEvery manifest input range processedA partial run can look plausible
For large outputs, aggregate checks narrow the search before detailed key comparisons. Compare partition-level counts and sums, then inspect mismatched keys. Hashes can help detect differences if serialization, ordering and null representation are defined, but a matching hash is not a substitute for choosing the right business columns and source range. Do not hash only keys when the suspected bug changes amounts.
Promotion needs a visibility contract. A pointer switch to a complete output generation can give readers a clean version boundary when the storage and query system support it. An in-place merge may expose a mix during the run. That can be acceptable for some dashboards and unacceptable for a financial close. State the reader behavior and verify the actual engine's atomicity rather than assuming the word “merge” guarantees a table-wide release boundary.
A safe stop leaves progress evidence. Record which input ranges completed and which output generation they belong to. Resuming with the same pinned transform should not duplicate results. If the transform must change halfway through, start a new generation or reprocess the affected range under a clear version rule. Do not combine old and new semantics while keeping one unqualified run label.
Resource limits protect the live pipeline. A replay can consume database connections, broker bandwidth, shuffle space and storage I/O even when it uses separate compute. Monitor the shared bottleneck and use a stop condition tied to live freshness or error rates. A low CPU percentage on the replay worker does not prove that the production sink is unaffected.
Finally, rerun a small complete partition into a second disposable generation and compare keyed output. This checks deterministic behavior under the same inputs and parameters. It does not prove the transformation is semantically correct; the expected-difference checks provide that separate evidence. The combination of deterministic replay, business validation and a reversible reader switch makes the repair reviewable.

Worked example

Teaching repair covers three dates. Old outputs contain 100, 120 and 80 unique orders. A corrected parser should add two orders on the second date and remove one invalid order on the third. Expected new counts are 100, 122 and 79, totaling 301 instead of 300. Build into a versioned table. Check counts by date and the three named changed keys, then switch readers after validation. A total of 301 alone is insufficient because the additions and removal could be assigned to the wrong dates.

Exercise

A replay creates counts 101, 121 and 79, total 301. Expected counts are 100, 122 and 79. Should it replace the old output? Give the next diagnostic step.

Model solution and rubric

No. The matching grand total hides a partition mismatch. Compare keys in the first two dates and inspect timestamp parsing, timezone conversion and partition assignment. Keep the old output available. Resume promotion only when expected differences are explained and business-key checks pass. Also verify that the replay did not trigger outward side effects or reduce live freshness beyond its limit.
Score out of four: one point for the correct result, one for showing the intermediate reasoning, one for identifying the stated failure case, and one for a verification that could disprove the answer. Do not award the reasoning point for a tool name alone.

Failure modes and misconceptions

“A separate consumer group isolates all replay effects.” It isolates offset ownership, but shared sinks, external calls and infrastructure can still be affected.
“The job exited successfully, so the repair is complete.” Completion requires every manifest range and the expected business result, not only a process exit code.

Interview probe

Evidence class: recommended. Original practice.
How do you know a backfill is safe to rerun?
Strong answer: Its input range and transform version are fixed, output keys are deterministic, repeated execution converges, and external effects are isolated or idempotent. I would verify a second run in a disposable destination and compare both keyed results.
Follow-up: What if the source retention is shorter than the required repair horizon?
Weak answer indicators: Validating only total row count; replaying into live tables without a visibility plan; using wall-clock now inside historical transforms.

Sources

Technical references: Airflow best practices; Kafka 4.1 design; dbt data tests. Sources support the documented mechanisms. The numbers, decisions, rubrics and interview prompts in this lesson are original teaching examples, not measurements or employer question claims.
docsAirflow best practicesairflow.apache.orgdocsKafka 4.1 designkafka.apache.orgdocsdbt data testsdocs.getdbt.com

Checkpoint

Expected counts are 100,122,79; replay produces 101,121,79. Grand totals match. Promotion?

AHold and inspect date/key differencesBPromote because 301 matchesCReplace expectations with observed countsDMerge all dates into one partition
Sign up free to answer and see why

Checkpoint

Historical eligibility uses the current clock, so reruns differ tomorrow. What makes the time-dependent rule reproducible?

APin only the source partitionsBPin an explicit logical as-of time with input and transform revisionCUse the job's actual finish time in every partitionDSeed the task scheduler but leave the clock lookup unchanged
Sign up free to answer and see why

Checkpoint

Replay writes to an isolated table, but each processed row also sends a notification. What must be addressed?

AIsolation of table readers automatically isolates notificationsBA new consumer group guarantees notifications are uniqueCSuppress or make external effects idempotent under the repair contractDUse faster writes so fewer notifications overlap
Sign up free to answer and see why

Checkpoint

The source range expired before repair. What is defensible?

ATreat the missing range as zeroBIncrease retention and claim recoveryCFill it from current state without labelling the differenceDFind another authoritative source or report the reconstruction gap
Sign up free to answer and see why

Checkpoint

Replay worker CPU stays low while live sink freshness deteriorates. What should be inspected next?

AShared database I/O, connection pools and storage limitsBOnly the replay's allocated CPU limitCOnly the live pipeline's source timestamp formatDOnly whether the replay and live job have different consumer groups
Sign up free to answer and see why

Explain how you would plan a deterministic replay with bounded impact without reading the solution. State one assumption that could change your answer, and one observation that would make you revise it.

Not yetGetting thereConfident

Wrap-up

  • A replay needs a fixed input, deterministic output and a promotion check. Keep the live workload and external effects under control.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.