Lesson 4 of 4 · 60 min

Defend the result, including its limits

Present a work sample with a reproducible result, known gaps, and a justified next step.

A work sample is evidence of how you solve a problem. The final discussion should let another engineer run it, inspect the important decision, and understand what you did not verify. A polished interface cannot compensate for a result whose correctness is unclear. Nor does a long list of future improvements explain why the present design fits the task.
Prepare a short execution path. State the input contract, how to run the sample, the expected output for a small fixture, and the most useful failure case. Show one complete result before discussing architecture. This gives the interviewer a reference point for later trade-offs. If part of the sample uses a fake dependency, label it clearly and explain the behavior that the fake does and does not represent.
Performance claims need a workload and a measurement method. A statement that an endpoint is fast is not falsifiable. A statement that a local fixture of 10,000 records completes in a measured time is more useful, provided you state hardware, repetitions, warm-up, and whether the number includes network and startup costs. Never invent a benchmark because the output looks responsive.
Use limitations to guide a next step. A local in-memory deduplication map may be appropriate for demonstrating logic in a short exercise, but it does not survive process restarts or coordinate multiple workers. Explain that exact boundary. A candidate who knows why a shortcut is limited gives the reviewer more confidence than one who calls every shortcut production ready.

Worked example

A fictional export worker submission includes fixture F1 with 12 input rows, two invalid-date rows and one separate duplicate-business-key row. These invalid and duplicate counts are disjoint. Its contract rejects invalid rows with reasons and retains one row for each valid business key. The demonstration produces nine accepted rows and three rejected records. A repeat run produces the same logical output, with a new execution timestamp kept outside the content comparison.
The candidate explains that the work sample uses local files and a single worker. The next production step is durable operation tracking and a storage-appropriate publication protocol. This fictional sample writes destination output in place before completion, so a crash can expose partial output; a verified atomic local rename would have different behavior. Adding a second worker before solving that issue would create more failure schedules.

Produce a fixture manifest

A work sample needs a known expected result. The following invented fixture makes data interpretation visible before performance or interface polish.
RowBusiness keyDateStatusExpected handling
1AValidOpenAccept
2BValidOpenAccept
3BValid, same valuesOpenDuplicate delivery
4CInvalidOpenReject with date reason
5DMissingOpenReject under stated contract
This smaller manifest is separate from the earlier twelve-row example. Its expected result is two accepted business records and three rejected or duplicate input records under the declared reporting categories. Define whether duplicate delivery belongs in a rejected count or a separate duplicate count before displaying totals. A sum that mixes categories without stating their meaning can look correct while misleading the user.
A completion record could be:
code
1fixtureId: F-small2inputRows: 53acceptedBusinessRecords: 24invalidRows: 25duplicateRows: 16outputKeys: [A, B]7contract: invalid dates rejected; identical duplicate keys collapse8unverified: conflicting duplicates, interruption, concurrent writers
The record states what passed and what remains. It does not imply that every production input is valid. A reviewer can change row 3 to conflict with row 2 and ask how the policy should respond. That is a useful discussion because it tests whether the candidate understands identity and ambiguity rather than only the happy path.

Make the run instructions verifiable

A good README tells the reviewer the required runtime, setup command, execution command, fixture location, expected output, and reset boundary. If a database or external account is not needed, say the sample runs locally. If a fake dependency is used, name it and state which behaviors it simulates. Do not require hidden credentials or an already-populated local cache.
Include one failure demonstration that is safe and deterministic. For example, a supplied invalid-date fixture can show error reporting without changing the user's environment. A crash schedule can be simulated by pausing before final result publication. The exercise should not require killing an unrelated process or deleting arbitrary directories.

Compare performance honestly

Suppose version A processes 10,000 rows in 900 milliseconds and version B processes the same rows in 450 milliseconds on the same machine under the same measurement boundary. That is a twofold improvement for this measured condition, subject to run variation. If B instead uses 1,000 rows, the comparison is confounded. A smaller input does not prove a faster algorithm.
Report several runs when making a benchmark claim and distinguish startup from steady-state work. Include whether input reading, validation, output writing, and network calls are inside the measurement. A database plan's execution time cannot stand in for a full user request if transfer and rendering dominate. A work sample can honestly omit a formal benchmark and simply state that performance has not been characterized.

Misconceptions to correct

The first misconception is that a passing fixture proves production readiness. A fixture checks a specified behavior for specified input. Concurrency, provider integration, permission boundaries, and recovery require their own evidence.
The second misconception is that a long future-work list demonstrates good judgment. A useful next step is the highest-impact unresolved risk under the current contract. Adding a new framework while the output can be partial after a crash is difficult to defend.

Extend the exercise

The reviewer gives one extra hour. The sample already computes correct results but can publish a partially written file if interrupted. Propose the next artifact. A strong answer adds a temporary-output and atomic-or-recoverable publication design, plus an interruption test. It states that storage-specific atomicity must be verified. Award one point for selecting the correctness risk, one for the publication boundary, and one for the test's final content assertion.
A final walkthrough can now show the fixture, result manifest, one failure path, and one decision record in a few minutes. These artifacts make the candidate's reasoning inspectable without requiring a large application.

Exercise and solution

A reviewer asks whether nine accepted rows proves the implementation is correct. The model answer says no. It proves the expected result for one fixture. Add cases for all-invalid input, duplicate keys with conflicting values, empty input, and interruption before result publication. Award one point for limiting the claim, one for two meaningful edge cases, and one for a failure-injection case tied to the implementation boundary.

Interview probe and wrap-up

What would you change with one extra hour? A strong answer chooses the largest unresolved correctness or usability risk and states a check that would confirm improvement. Follow up by removing that hour. A weak answer lists caching, microservices, and a new framework without linking them to a measured problem. Make your work easy to run, your claims easy to check, and your shortcuts easy to find.

Sources

docsMonzo backend interview guidance, March 2025monzo.comdocsPostHog engineering work-sample guidanceposthog.comdocsPostgreSQL EXPLAIN and execution analysispostgresql.org

Checkpoint

The five-row fixture has A, B, duplicate B, invalid C, and missing-date D. Under the stated contract, accepted business records?

AFive.BOne.CTwo.DThree.
Sign up free to answer and see why

Checkpoint

Version B uses one-tenth the input and runs twice as fast. What is supported?

AThe new implementation is slower per row in every case.BA twofold algorithm speedup is proven.CProduction throughput doubled.DThe runs are not a comparable speed benchmark.
Sign up free to answer and see why

Checkpoint

A fake provider returns success deterministically. What remains unverified?

AOnly the sample's language syntax, because any provider with the same response shape behaves identically.BReal authentication, network behavior, and provider recovery semantics.CThe fixture's expected local result only.DOnly the local fixture's arithmetic, because the fake makes external behavior equivalent.
Sign up free to answer and see why

Checkpoint

The sample can expose partial output after interruption. Best next-hour priority?

AAdd a recoverable publication protocol and interruption test.BExtend the worker lease while the destination remains visible during writes.CRetry interrupted work by appending to the same visible output file.DMark the job failed after partially written output becomes visible.
Sign up free to answer and see why

Checkpoint

What belongs in a defensible completion note?

AThe successful local run described as proof of all production conditions.BA future-work list without the current fixture result.CA list of passing tests without their input or boundary definitions.DExact verified fixture outcomes and explicit unverified boundaries.
Sign up free to answer and see why

Can you present a fixture and run contract, limit your claims to the evidence, and select the next hour of work by the largest unresolved user risk? State the relevant identifiers, failure boundary, and evidence in your own words before selecting your confidence.

Not yetGetting thereConfident

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.

Defend the result, including its limits · Backend…