Treat progress messages as hints about durable state
Recover from missing, duplicated, and out-of-order progress events.
Streaming progress makes a long task feel understandable. It also introduces another delivery path that can disconnect or replay. The durable job record remains the source of truth. A progress event tells the client that a state may have changed; it should not be the only place where completion exists.
Server-sent events provide a one-way stream from server to browser. The event format supports identifiers, and clients can reconnect. That does not automatically implement durable replay for your application. The server must decide whether it retains events, how it interprets a last-seen identifier, and what happens when the requested history has expired. A status fetch is often the simplest recovery path for a job interface.
Give job states monotonic versions. If the client has version 7 and receives version 6, it should not move backward. A duplicate version can be ignored if its content is consistent. A gap from version 4 to version 7 can trigger a current-state read instead of requiring the client to reconstruct every intermediate percentage. Progress and terminal state need different handling: missing a 40 percent update is usually acceptable; missing the only record of completion is not.
Clean up subscriptions when the visible job changes or the component unmounts. Otherwise an old stream can continue updating a new page or consume connections. Cancellation of the subscription is separate from cancellation of the job. Closing the tab does not necessarily cancel work, and the interface must not imply that it does.
Worked example
A fictional client displays J8 at version 3, running at 30 percent. It receives:
The client accepts v5, ignores stale v4 and duplicate v5, then fetches status after reconnect. It displays completion from v6. It does not average percentages or infer failure from the lost stream. A new subscription can continue from a documented replay position if the server supports one.
Separate transport position from job version
A stream event ID and a job version solve different problems. An event ID can be a cursor in a replay log. A job version orders authoritative states for one job. If a stream carries several jobs, the cursor may increase globally while each job has its own version. Comparing them as the same counter can discard a valid update or accept an old one.
Record a feed cursor under the stream's replay contract and J8 version 6 under J8's state model. If the server restarts feed numbering, it needs an epoch or another unambiguous cursor definition. The presence of an id line in an SSE message does not establish durable replay.
A complete snapshot can replace earlier state when it has a newer valid version. A delta cannot necessarily do so. Suppose version 4 adds five processed rows and version 6 adds three. Missing version 5 means the client cannot derive the total from received deltas. A status snapshot reporting total processed rows repairs the view. State whether events are snapshots, deltas, or invalidation hints.
Recover without inventing history
Observation
Client action
Reason
Same job, lower version
Keep current state
Delivery can be stale
Same version and content
Ignore duplicate
No new state
Same version, conflicting terminal content
Fetch authority and record anomaly
Version contract violated
Gap with complete snapshot
Accept valid newer snapshot
Intermediate views may be unnecessary
Gap with deltas
Fetch snapshot or complete replay
Missing changes prevent reconstruction
Version checks are necessary but insufficient. A larger number for the wrong job must not enter the current page. Check job and account identity. When the user changes accounts, close the old subscription and reject obsolete callbacks. Cleanup reduces work; identity gating protects state when cleanup races with delivery.
Terminal snapshots also need transition validation. If the client has succeeded v9 and receives running v10, the expected state machine may reject the transition despite the larger number. Some products permit a new attempt under a new job revision, but this must be explicit. Counters should describe valid transitions rather than authorize arbitrary regressions.
Choose polling through a budget
Take an invented workload of 2,400 open pages. Polling each every six seconds produces about 400 status requests per second before retries. At thirty seconds, it produces about 80 per second. These calculations assume one request per page per interval and ignore jitter, background suspension, and failures. They are estimates from supplied assumptions, not product measurements.
A thirty-second interval can delay display of completion by almost thirty seconds after a poll, plus request time. If the product accepts that delay and jobs change state only four times, polling may suffice. A stream can improve freshness and reduce repeated reads, but adds connection and reconnect handling. Persistent connections have operational costs too.
For either method, use bounded retry behavior consistent with the API. Jitter helps prevent many clients from reconnecting together. A hidden tab can reduce polling under the product's policy; returning to the tab should fetch current state. Stop unnecessary polling after a verified terminal state while retaining a refresh path for result access changes.
A status service also needs permission checks during recovery. Do not replay cached events from a previous account merely because the cursor is recent. The reconnect request establishes the current authorized context. If permission has expired, show that condition without revealing the private terminal result. Transport reconnection cannot restore revoked authority.
Misconceptions and a second exercise
One misconception is that automatic browser reconnect guarantees every application event is replayed. Replay depends on retention and cursor interpretation. A second is that a percentage is intrinsically durable. A worker can discover more rows or restart, so define whether progress may reset or reflects a stable fraction.
Exercise: a client holds J8 version 4 from a delta stream. Its cursor is older than server retention. The server reports history unavailable, while authorized status reports succeeded version 8. Accept the terminal snapshot, reset the cursor under the protocol, and do not rerun J8. Award one point for distinguishing cursor and version, one for snapshot recovery, one for retaining job identity, and one for stopping unnecessary progress polling.
In an interview, show that completion remains retrievable without any stream connection. Then explain whether live updates help the user's next action or merely animate waiting. Reliable progress communicates known state and uncertainty; a smoothly increasing display alone proves neither.
Exercise and solution
The client receives a succeeded event for J9 while its page shows J8. The handler shares one global progress variable. Identify the repair. Key subscriptions and stored updates by job ID, check resource identity before changing visible state, and close the old subscription on navigation. Award one point for each. Merely adding a delay before rendering does not fix the wrong-resource update.
Interview probe and wrap-up
When would polling be preferable to a stream? A strong answer considers update frequency, operational complexity, proxy behavior, client count, and required freshness. Follow up with a five-minute job whose state changes four times. A weak answer chooses WebSockets because real-time sounds advanced. Choose the delivery method after defining recovery. The user needs accurate status more than every intermediate event.
A client misses a delta, then receives a later delta. Safe recovery?
AAdd only the received deltas and label the total authoritative.BFetch a snapshot or complete documented replay.CStart a replacement job.DAssume the missing change was zero.
The same job version has conflicting terminal states. Best response?
AKeep the last arrival silently.BFetch authority and flag the contract violation.CPick the state with the later worker clock timestamp.DChoose succeeded as the useful outcome.
Can you distinguish replay cursor from job version and recover missing deltas without repeating the job? State the relevant identifiers, failure boundary, and evidence in your own words before selecting your confidence.