Write a rollout and rollback record that accounts for stored data.
A release changes both user behavior and system state. A plan that says deploy and monitor is incomplete because it does not say what evidence should stop the release or whether the previous version can safely run afterward. Define the exposure group, success condition, failure threshold, and rollback boundary before the change reaches users.
Separate deployment from release. Code can be present while a feature remains disabled for ordinary users. A limited release can expose it to a named pilot group. This reduces the number of affected users, but it does not remove the need for authorization or data compatibility. A feature flag controls whether a path is used; it does not automatically reverse data that path already wrote.
Data changes need a compatibility plan. Adding a nullable field is often easier to roll back than replacing one field with another and immediately deleting the old value. An expand-and-contract sequence first supports both representations, then migrates or backfills data, then changes reads, and only later removes old support. The exact sequence depends on the application and requires verification rather than a generic recipe.
A stop condition should relate to the user outcome. An export rollout might stop when valid jobs fail above a threshold, completion latency exceeds the pilot target, or outputs differ from expected totals. CPU can help explain a failure, but a quiet CPU chart does not prove correct exports. Keep the observation window long enough to include the relevant workload.
Write each comparison rule before seeing results. Exact customer membership and integer minor-unit totals may require exact equality. A different contract may permit a stated absolute error for an unrounded calculation in a fixed currency. For example, if that separate contract allows absolute error of 0.01, expected 100.000 and observed 100.004 pass the numeric comparison because their difference is 0.004. This does not waive membership, rounding, access, or latency checks. Do not apply that tolerance to the integer migration below, whose conversion must preserve exact minor units. A failed check and an allowed numeric difference are distinct outcomes.
Worked example
A fictional release changes report totals from a stored integer field total to a structured amount/currency representation. Version A reads total. Version B can read both and writes both consistently. The team deploys B, verifies old records, then backfills amount/currency in bounded batches. A later version reads the new representation only after comparison checks pass. Removing total waits until the rollback window closes.
The release record says pilot accounts only, compare twenty representative exports, stop on any unexplained total mismatch, and preserve the old field until compatibility review. If B fails early, returning to A is possible because A still has the data it understands.
Define the representation before migrating it
The example needs a precise conversion contract. Assume the old total is an integer in USD minor units, so 1234 means 12.34 USD. The new representation is amountMinor=1234 and currency=USD. Do not infer currency from a customer's location or divide by 100 into an imprecise floating-point value without a defined contract. If old rows actually mix currencies without metadata, backfill cannot reconstruct missing information by guesswork.
The migration has separate reader, writer, and data states:
Phase
Readers
Writers
Safe return to A?
Before expansion
Old total
Old total
Yes
B deployed
Both representations
Both consistently
Yes, if every new row keeps total
Backfill verified
Both representations
Both consistently
Yes within stated compatibility
New-only reads
New fields
Compatibility writes retained
Depends on verified old-reader support
Old field removed
New fields
New fields
No simple code-only return
This is a teaching sequence, not a universal migration script. Locks, schema size, deployment order, and actual database behavior need separate verification. The important idea is that compatibility belongs to the combination of running code and stored data.
Verify bounded backfill progress
Use a resumable batch marker and comparison counts. Do not report migration complete from one successful batch. A record might look like this:
code
1migration: report-money-v22range: report IDs1000..14993examined:5004alreadyCompatible:1205converted:3796heldForReview:17comparisonMismatch:0 among499 compatible rows8heldReason: legacy total missing9nextRange:1500..1999
The held row is not silently converted to zero. Its missing value is a data-quality problem requiring a defined repair. Count it separately so the sum reconciles. Verify new writes during the backfill too: otherwise all historical rows can look correct while the live writer creates incompatible records.
A concurrent update can race with the backfill. If the batch reads total, another writer changes it, and the batch later writes a stale new representation, the two fields diverge. Use a guarded update or a transactional protocol appropriate to the actual schema. A teaching pseudocode condition is update new fields from total only while the expected legacy version still matches. On mismatch, reread and retry under the migration contract.
Choose a stop signal and a stop action
The twenty-export comparison includes old records, new records, zero values, large values, and missing-data handling under the defined policy. A mismatch should stop expansion of exposure, preserve the failing record's identity and versions, and trigger investigation. It should not automatically erase the new field or rerun every batch without understanding the cause.
Separate three actions: disable new feature exposure, stop further backfill batches, and change running code. They can have different safety consequences. If B still writes both representations correctly but the new display has a formatting defect, disabling the new display may be enough. If B corrupts both fields, the team needs immediate containment of the writer and a verified repair plan.
A quiet error dashboard is not proof of correct conversion. The release could produce numerically plausible wrong totals without exceptions. Comparison fixtures detect semantic errors that process health cannot. Likewise, a small pilot provides limited load evidence even when every fixture passes.
Misconceptions and a second exercise
One misconception is that dual write means two unrelated writes can never diverge. Consistency requires an atomic or recoverable protocol and shared conversion rules. Another is that the previous binary is always a safe fallback. Its data assumptions may no longer hold after contraction.
Exercise: batch 1499 finishes, but one row was held for missing total. New code reads only the new representation. Should the rollout declare all rows migrated? No. Record the unresolved row and either prevent that reader path for the affected record or repair it under an approved data policy. Do not manufacture zero. Award one point for reconciled counts, one for preserving the held case, one for reader safety, and one for a concrete verification before resuming.
The release packet should let another engineer identify current phase, exposed accounts, last verified batch, stop controls, and compatible rollback versions. This is more useful than a single checkbox labelled migration done. It supports safe continuation when the original author is unavailable.
Exercise and solution
The team already deleted total, and version A crashes when reading new records. Is code rollback alone safe? No. The model answer identifies the data compatibility failure and requires a verified data recovery or forward-fix plan before changing versions. Award one point for distinguishing code from data, one for preserving evidence, and one for avoiding an untested destructive reversal.
Interview probe and wrap-up
What would you do if the pilot has too little traffic to evaluate reliability? A strong answer uses deterministic fixtures and controlled smoke tests while stating that real-load evidence remains limited. Follow up with a failure affecting only old records. A weak answer declares success because no alert fired. A release is reviewable when another operator can identify who is exposed, what to measure, and what recovery still remains possible.
A backfill reads total then writes after another update. What prevents stale conversion?
AA version-guarded or transactional conversion protocol.BSort the batch by ID but leave the read and update unguarded.CRun the backfill only once, while ordinary writers remain active.DCheck that the new field is non-null, without comparing the source version.
One of 500 rows is held for missing total. What can be claimed?
AAll 500 were converted correctly.BThe missing total is necessarily zero.CThe reader can ignore the held record forever.D499 compatible rows checked; one unresolved case needs a safe policy.
Old fields were removed. What blocks simple code rollback?
AKeeping old source code guarantees the old fields still exist.BOld readers may no longer understand stored records.CThe old binary compiled successfully before the migration.DDisabling the new UI automatically restores the old data shape.
A comparison finds wrong totals without exceptions. Best stop criterion?
ADelete all new fields immediately.BContinue because process health is green.CStop exposure expansion and investigate the mismatch with record/version evidence.DIgnore the fixture because it is small.
Can you explain reader/writer/data compatibility at each phase and stop a rollout without assuming old code can read new data? State the relevant identifiers, failure boundary, and evidence in your own words before selecting your confidence.