Lesson 1 of 4 · 25 min

Join the snapshot to the change log

Explain a gap-free CDC snapshot and streaming handoff.

Mechanism and reasoning

Change data capture reads database changes and turns them into a stream. An initial snapshot supplies existing rows; the log supplies later changes. The hard part is the boundary. If the snapshot and log positions do not align, rows can be missed or changes can be applied twice. Use the connector's documented snapshot protocol rather than inventing a timestamp cutoff.
A snapshot is not necessarily a single instant unless the database isolation and connector process provide that view. Changes can continue while rows are read. The connector must record a log position and coordinate the snapshot with subsequent streaming. The sink still needs the identity and ordering rules from the foundations course.
Deletes require explicit treatment. Removing a row from the source should not leave it active forever in a downstream current-state table. Preserve the operation type and key. Some stream formats also emit tombstones for compaction; distinguish a business delete event from transport cleanup behavior.
The source log has operational cost. A replication slot or equivalent retention mechanism can hold log segments while the consumer is behind. A stalled connector can therefore threaten the source database's disk even if the analytical destination is separate. Monitor lag in bytes and time, source disk and connector health. Set an escalation policy before storage is exhausted.
Recovery may require a new snapshot if the necessary log range is no longer available. This changes load and can duplicate records unless the sink converges. In an interview, explain the successful path and the source-protection plan. 'We use CDC' is incomplete without handoff, deletion and lag behavior.

Review the handoff as a consistency protocol

A snapshot and a change stream describe overlapping views of the source. The connector's job is to combine them without creating an unobserved interval. The exact mechanism depends on database isolation, log positions and connector mode. Read the documented protocol for the selected connector version. Do not replace it with a home-made “snapshot completed at noon, so read changes after noon” rule, because wall-clock time is not the source log's consistency boundary.
code
1Illustrative unsafe handoff209:00 snapshot starts reading customer table309:03 customer c1 changes after its row was read409:04 customer c2 is deleted before its row is read509:10 snapshot finishes609:10 naive consumer starts reading only new changes7Result: c1's update and c2's deletion event can be missed
The example does not reproduce Debezium's actual snapshot algorithm. It demonstrates why the connector must coordinate a consistent view and a resumable log position. With a supported connector protocol, overlap can be handled using its event metadata and the sink's duplicate-safe ordering rule. A timestamp cutoff alone cannot establish that every change is represented.
Source identity also matters. A primary-key change can be represented through delete and create semantics under a connector's documented event format. A downstream table that treats it as an ordinary update to an immutable key may leave the old key active. Test the source operations that the application actually uses: insert, update, delete, key change and schema change. Do not infer all their representations from one successful insert.
Operational signalWhat it can revealWhat it does not prove
Retained log bytesSource storage pressureDestination business correctness
Source position minus consumed positionCapture lag under defined unitsTime until a specific report is valid
Oldest unprocessed event ageFreshness delayComplete source coverage
Source free diskRemaining storage headroomSafe automatic deletion of logs
Sink accepted versionApplied current-state progressCorrect handling of every entity
Lag has multiple stages. The connector can capture changes promptly while the broker consumer is slow. The broker can be current while an analytical transformation is delayed. A single “CDC lag” number can hide which component needs action. Track the source capture position, transport progress and sink publication position separately, with a clear definition for each metric.
Estimate time to exhaustion conservatively. If retained logs grow at 2 GiB per hour and 8 GiB remain, four hours is an upper bound under the simplified constant-rate model, not an alert time. Other source writes, bursts and operational response time reduce the safe margin. An intervention threshold should leave time to diagnose and act. The lesson's calculation excludes those effects deliberately so the arithmetic is inspectable.
A source-protection action can sacrifice downstream continuity. Dropping a replication slot or otherwise releasing retained logs may protect the database but remove the resume point. That can be the right emergency choice under an approved incident policy, yet it must be recorded as a recovery tradeoff. The subsequent sink may need a fresh consistent snapshot and reconciliation. Do not silently continue from the newest log position and claim that no changes were lost.
A resnapshot also consumes source resources. Limit concurrency, understand isolation and locking behavior, and measure read pressure. Rebuilding into a separate sink generation can avoid exposing partial current-state output. The duplicate and deletion rules remain necessary because source state can change during the process. A successful connector restart is only one event in a longer recovery sequence.
For a practical drill, use a small disposable source and perform updates and deletes while an initial snapshot runs. Stop the connector, generate enough changes to create measurable lag, resume within the retained range and compare the keyed sink against the source. Then test the documented behavior when the range is unavailable. The exercise establishes which recovery paths work before a real source-disk incident forces a hurried choice.
The missed deletion event has a different effect depending on sink state. A clean, initially empty latest-state sink that never received c2 may already omit it correctly, even though its event history is incomplete. A reused sink that still holds c2 can leave a stale row unless the deletion or an explicit full-snapshot reconciliation removes it. Do not claim that every missed deletion necessarily creates a wrong current row; name whether the contract is event completeness, current-state convergence, or both.

Worked example

Invented source has customer c1 at version ten when the coordinated snapshot begins at log position P100. During the snapshot, c1 changes to eleven at P105 and c2 is deleted at P108. The connector emits its snapshot records and the relevant changes using its documented handoff. The sink compares versions or ordered source positions and ends with c1 at eleven and c2 absent. A naive filter that starts streaming only after the snapshot finishes can miss both changes.

Exercise

A connector is stopped for six hours. The source generates 2 GiB of retained log each hour and had 8 GiB free at stop time. Ignore other disk use. When does this become unsafe, and what must recovery check?

Model solution and rubric

The retained log consumes eight GiB after four hours, so waiting six hours exceeds the available space under the simplified assumptions. Alert and intervene before that point. Recovery must check whether the required log positions still exist and whether the connector can resume. If not, use a controlled new snapshot and duplicate-safe sink. Do not discard source logs blindly to make space, because that can remove the recovery path.
Score out of four: one point for the correct result, one for showing the intermediate reasoning, one for identifying the stated failure case, and one for a verification that could disprove the answer. Do not award the reasoning point for a tool name alone.

Failure modes and misconceptions

“Snapshot finish time is a safe streaming start point.” Changes can occur after a row is read but before the snapshot ends. Use the documented consistency boundary.
“The analytical pipeline cannot harm its source.” Retained logs and resnapshot reads can consume source disk and capacity. Monitor and limit both.

Interview probe

Evidence class: recommended. Original practice.
Why can a broken analytics connector take down the transactional database?
Strong answer: Its replication state can retain source logs and consume disk. I would monitor retained bytes and lag, define an intervention threshold and verify a resumable log range before restarting or resnapshotting.
Follow-up: How would a primary-key change appear in your connector's event contract?
Weak answer indicators: Using snapshot finish time as a log boundary; ignoring deletes; treating CDC as zero-load on the source.

Sources

Technical references: Debezium PostgreSQL connector; Kafka 4.1 design. Sources support the documented mechanisms. The numbers, decisions, rubrics and interview prompts in this lesson are original teaching examples, not measurements or employer question claims.
docsDebezium PostgreSQL connectordebezium.iodocsKafka 4.1 designkafka.apache.org

Checkpoint

A row changes after its snapshot read but before snapshot completion. Which handoff is unsafe?

AThe connector's documented log-position protocolBA duplicate-safe sink using source orderCStarting only from wall-clock snapshot completion timeDA consistent snapshot with coordinated streaming
Sign up free to answer and see why

Checkpoint

Retained logs grow 2 GiB/hour with 6 GiB free. Ignoring other use, exhaustion occurs after?

A12 hoursB4 hoursC2 hoursD3 hours
Sign up free to answer and see why

Checkpoint

Why should the alert occur before that theoretical exhaustion time?

AOther writes, bursts and response delay consume marginBThe formula proves logs stop growing automaticallyCDestination row counts eliminate source riskDThe snapshot always releases every retained log
Sign up free to answer and see why

Checkpoint

Source capture is current but analytical publication is hours behind. What should be inspected next?

ACapture lag alone, because it represents every downstream stageBOnly whether the source accepts writesCThe broker-consumer, transformation and publication progress boundariesDOnly the source's replication slot name
Sign up free to answer and see why

Checkpoint

Required log positions are unavailable after source protection. What is defensible?

ASilently start at the newest offsetBTreat missing changes as no-opsCUse the documented recovery/resnapshot path and reconcile the sinkDAssume offset names reconstruct the lost changes
Sign up free to answer and see why

Explain how you would explain a gap-free cdc snapshot and streaming handoff without reading the solution. State one assumption that could change your answer, and one observation that would make you revise it.

Not yetGetting thereConfident

Wrap-up

  • The snapshot/log boundary and source log retention are part of the pipeline contract. Design both before the first full load.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.