Trace the complete failure, then test that schedule
Write an end-to-end test plan using a deterministic failure trace.
A test that checks whether the export button appears does not prove that an export survives a lost response. The important behavior spans browser, API, worker, and result storage. Select a small number of failure schedules that exercise the contract at those boundaries. Each test needs a controlled trigger and observable final state.
Use deterministic substitutes where they preserve the behavior under test. A fake worker can pause after claiming a job, complete it on command, or return a permanent validation error. A fake clock can advance retry delays. These tools make race schedules repeatable. They do not prove the actual provider's network behavior or storage permissions, so retain a separate integration check for those boundaries.
Trace identifiers connect observations but must not become business identities. One operation can have several request traces and worker attempts. Record the operation key, job ID, attempt ID, state version, and result ID in structured events. Avoid logging complete input files or credentials. The trace should answer what happened without becoming a second store of private data.
Tests should assert the user-visible truth and the durable effect. After a lost response, a correct interface may show pending confirmation while the job continues. The final test should confirm that exactly one logical job exists and that the user can retrieve its result. Counting button clicks or network calls alone is weaker evidence.
Worked example
Use an invented two-row fixture with values 10 and 20. The export should contain two rows and total 30. Pause the API after it commits J8 but before replying. Drop the reply. Let the client retry with K8, then release the worker. The resulting trace is:
Assert one job for K8, one complete result, correct two-row content, and visible success after recovery. The system may contain two request attempts. That is expected, so a test requiring one HTTP request would reject valid recovery.
Turn the test into a causal experiment
A useful failure test controls the boundary that could violate the guarantee. Write setup, pause point, competing action, release condition, and final assertions before selecting a framework. This lets a reviewer understand the test without knowing the harness.
Here is a second invented schedule for private-output publication:
code
11. Submit K8; persist J8 and dispatch intention.22. A1 prepares private O1; pause before pointer binding.33. Expire A1's lease under the defined clock.44. A2 claims J8, prepares O2, and binds O2 with succeeded.55. Resume A1 and attempt its old guarded completion.66. Read status and content through the authorized product route.
J8 must remain succeeded with A2's permitted pointer. A1's transition affects zero rows or returns a stale-owner outcome. O1 may remain private pending cleanup. The test should not fail merely because two private objects exist. It should fail if O1 becomes visible or replaces the winning result.
A fake storage adapter can make immutable writes and pointer checks inspectable. It cannot prove that the real store blocks public access to unbound objects. A separate integration check must establish storage policy, signed-access behavior, and content retrieval. The local experiment tests state logic, not a provider configuration it never exercised.
Connect assertions to promises
Guarantee
Controlled trigger
Durable assertion
User assertion
One logical submission
Drop accepted reply
One job for scoped key
Same job recovered
No stale publication
Resume old attempt
Winning pointer unchanged
Correct content
Cancellation truth
Cancel before binding
No published pointer
Cancelling/cancelled
Stream independence
Disconnect before completion
Terminal state retained
Refresh finds completion
Access isolation
Read as another account
No permitted result access
No private content
This matrix exposes missing coverage. A disappearing spinner says little about one logical submission: a second job might finish with the same visual result. A database count says little about whether a user can recover the result. Full-stack verification needs durable effects and the corresponding user truth.
Use known content. Rows 3, 7, and 11 produce total 21 and three rows under the supplied arithmetic contract. If you add an invalid row, specify whether the contract rejects that row or fails the whole job. Without this rule, implementations with different behavior can each claim success.
Trace causality without copying payloads
A useful event contains operation ID, job ID, attempt ID, state version, and transition result, with appropriate account-safe correlation. A transport trace identifies one HTTP attempt. A worker attempt identifies one execution claim. Neither replaces the stable operation key across retries.
For rejected completion, log that A1 attempted old authority while A2 held the current claim. That explains the race without copying private report rows. A log emitted before database commit does not prove durable success. Wall-clock timestamps across machines also do not establish exact causal order; versions, relationships, and controlled barriers are stronger evidence.
You can verify the trace by walking from user intent to result pointer. Every HTTP retry should map to the same logical job under the replay contract. Every worker attempt should name that job and its own claim. Every exposed result should be reachable from the committed job pointer. An unexplained new job or pointer is a concrete investigation lead.
Permit valid variation
Two transport requests are allowed when the first response is lost. Two attempts are allowed after lease expiry. A private orphan is allowed after a crash. The required condition is that these do not create another logical job or expose an invalid result.
Do not accept any terminal state merely to reduce flaky tests. If cancellation demonstrably committed before guarded publication, success violates this protocol. If the harness cannot establish the winner, improve the barrier or state the weaker conclusion honestly. Longer sleeps do not create evidence of transaction order.
Misconceptions and a second exercise
One misconception is that timed waits make a race deterministic. Slow machines and scheduling variation break that assumption. Wait for observable barriers such as job committed, output prepared, or cancel transaction complete. A second is that every duplicate attempt is a correctness failure. Attempts can repeat safely; business effects and authorized visibility define the guarantee.
Exercise: a test uses a fake that immediately marks jobs succeeded and skips pointer binding. What claim is unsupported? It cannot test cancellation-versus-publication ordering or private-output visibility because those boundaries were removed. Replace it with a controllable stateful fake or a targeted integration harness. Award one point for the omitted behavior, one for replacement control, and one for an assertion about actual content.
Present one successful run manifest and one forced failure schedule in the interview. State what each proves and which external boundary remains unverified. This shows judgment without requiring a large application or a test count that hides weak assertions.
Exercise and solution
Now pause after F8 is prepared privately but before pointer binding. What must recovery establish? Which attempt still has authority to bind a complete result. The model solution reads durable job state, allows only the guarded current-owner transition to expose content, and leaves unbound attempt outputs private for cleanup. Award one point for the identified gap, one for stable identity, and one for checking actual content rather than just a success label.
Interview probe and wrap-up
Which single test would you add first with limited time? A strong answer selects the failure most likely to violate the core contract and explains why a unit test alone misses it. Follow up with what the test cannot establish. A weak answer maximizes test count without considering boundaries. Keep the fixture small and the failure schedule precise so a reviewer can understand the result.
The test leaves private O1 while O2 is committed. Is O1's existence alone a failure?
AYes, only one object may ever exist.BYes, computation retries are forbidden.CNo, private orphans are allowed under the visibility/cleanup protocol.DNo, so O1 may later replace the pointer.
Which barrier establishes cancellation-before-publication?
AWait until the cancel request is sent.BObserve a log emitted before the cancellation transaction commits.CWait 500 milliseconds after clicking cancel.DObserve cancel transaction commit before releasing binding.
What does a fake skipping pointer binding fail to test?
AOnly transport reconnect behavior; the fake proves publication correctness.BThe publication/cancellation boundary removed by the fake.COnly content formatting; a succeeded status proves result visibility is safe.DOnly retry timing; checking final job count proves the publication winner.
AOne scoped job, correct authorized content, and recoverable user status.BExactly one HTTP request and no retried private writes.COnly a success toast.DOnly an accepted log message.
AOperation IDs replace account checks.BTrace IDs are not useful.CEach trace should create a new business job.DSeveral transport attempts can implement one durable intent.
Can you use controlled barriers to prove one logical operation, valid publication, and user recovery while allowing safe retries? State the relevant identifiers, failure boundary, and evidence in your own words before selecting your confidence.