Lesson 3 of 4 · 60 min

Explain the waterfall with numbers

Choose a performance change using a browser-to-server timing breakdown.

Full-stack performance includes more than database execution. A user waits for navigation, network transfer, server work, JavaScript execution, and rendering. Several individually fast operations can form a slow serial chain. A useful interview answer identifies the critical path and measures the part that blocks the user's next action.
Distinguish total work from elapsed time. Two independent requests taking 200 milliseconds each can complete in roughly 200 milliseconds when started together, plus overhead. If one waits for the other without a real dependency, elapsed time becomes roughly 400. Parallelism is useful only when the requests are independent and downstream capacity supports it. Launching many requests simultaneously can create a new bottleneck.
A server response can also be quick while the page remains unusable. A large payload may take time to parse. A long main-thread task may delay input. Rendering thousands of rows when the user sees twenty creates unnecessary work. A loading skeleton changes perception but does not reduce the underlying delay. State whether an improvement changes actual completion time, time to useful content, or only visual feedback.
Caching needs a correctness budget. A public list that changes daily differs from a private permission-sensitive result. Specify freshness, key dimensions, and invalidation before making a cache recommendation. Do not hide a wrong query behind a long-lived cache and call the system fixed.

Worked example

A fictional page does this:
code
1load session:          100 ms2then load project:     250 ms3then load activity:    300 ms4render usable view:     80 ms5total:                 730 ms
After authentication, project and activity depend only on the authorized project ID, not on each other's responses. Starting them together changes the estimate to 100 plus max 250,300 plus 80, or 480 milliseconds. This saves 250 milliseconds under the supplied assumptions. It does not make server work free, and it must preserve authorization for both requests.
If activity actually requires a version returned by the project request, the proposed parallelism is invalid. The correct answer changes when the dependency changes.

Draw dependencies, then calculate

A timing table can hide dependencies if every row is simply added. Draw a graph with a start condition and the action the user is waiting to perform. The first worked example assumes both project and activity requests can start after authentication and a known project ID. It does not assume that every request can start at navigation time.
code
1navigation2  -> session1003      -> project250 --|4      -> activity300 -|-> render80 -> usable5critical path: session100 + activity300 + render80 =4806total request work:100+250+300 =650
The 480-versus-730 comparison is not a benchmark until measured. Here 480 is an estimate under the declared timings and zero added concurrency overhead. Actual parallel requests may contend for a connection, CPU, database pool, or rate limit. Measure again after the change rather than presenting arithmetic as an observed speedup.
Promise.all is one way to await independent promises together. It does not itself make sequentially awaited work concurrent, and rejection does not automatically cancel other work. Start independent operations under the correct authorization context, decide whether one failure should fail the whole view, and handle partial results only when the product permits them.

Measure a second bottleneck

Consider this invented capture for a large report page:
StageCold run msWarm run ms
Server response and transfer22090
Parse and transform480475
Render to usable interaction410405
Total within this boundary1110970
The warm cache reduces the first stage by 130 ms but leaves client work almost unchanged. A larger server cache may not address the dominant remaining cost. Bound the payload, reduce unnecessary transformation, or limit visible rows, then measure the same user action again. Do not simply quote the improved server duration as if total usability improved by the same percentage.
Suppose a bounded page reduces parse/transform to 70 and render to 85 while response/transfer is 100. The new estimate is 255 ms. The feature now shows fewer rows initially, so preserve access to the complete dataset through defined pagination or export. Verify sort and totals across the whole intended population; calculating a global total from one page changes correctness.

Choose a measurement boundary

Define whether time begins at navigation, button activation, or request dispatch. Define whether it ends at response arrival, useful content, or successful next interaction. These are different metrics. A query plan reports database work and cannot stand in for a full browser path. A screenshot of a loading skeleton proves visible feedback, not completed data.
Use repeated runs and report spread where possible. Compare the same fixture, device class, connection assumptions, cache state, and measurement boundary. A cold first load and a warm in-app navigation are both useful, but they answer different questions. Keep them separate instead of selecting whichever number favors the patch.
A small work sample can report an honest local measurement with limitations. It need not imitate a production percentile dashboard. State that the measured environment is local and that concurrency, real network behavior, and representative devices remain unverified if they were not tested.

Avoid optimizing a false dependency

An activity endpoint may require the project's current version to ensure consistent content. Parallelizing it before that version is known can produce a fast but mismatched page. The correct response could be a combined server endpoint, a shared snapshot contract, or accepting explicitly independent freshness. The design choice must follow the consistency requirement.
Similarly, prefetching private data before permission checks can cross a security boundary. It is not enough that the UI later hides the result. Preserve access rules when changing call order. Performance work must maintain the original user contract.

Misconceptions and a transfer exercise

One misconception is that total server work equals elapsed user delay. Overlap changes elapsed time without eliminating work. Another is that the largest component in a timing table must always be optimized first. It may not be on the critical path, or reducing it may violate consistency. Locate the blocking dependency first.
Exercise: after 100 ms authentication, project takes 250 ms. Activity takes 300 ms but needs project's version. A permissions-independent public help request takes 200 ms and can start immediately. Rendering takes 80 ms after project and activity. What is completion time if help also must be present? The project/activity chain is 730 ms total; help ends at 200 ms and does not extend it. Parallelizing help does not remove the project-to-activity dependency. Award one point for the dependency chain, one for 730, one for identifying help off the critical path, and one for rejecting invalid activity parallelism.
Present both the expected benefit and the verification plan. A strong interview answer says which dependency changes, what guarantee is preserved, and what comparable measurement will establish the result.

Exercise and solution

The server responds in 120 milliseconds, but parsing a large payload takes 500 and rendering takes 400. Which experiment comes first? Bound the returned data and visible work, then measure parsing and rendering again. Award one point for identifying the client cost, one for a bounded-data change, and one for measuring time to useful interaction. A database replica does not directly address the supplied evidence.

Interview probe and wrap-up

How do you prove a performance improvement? A strong answer compares the same workload and measurement boundary, includes variation, and verifies that correctness did not change. Follow up with warm versus cold caches. A weak answer cites a framework benchmark unrelated to this page. Show the critical path, the dependency you can remove, and the measurement that confirms the user benefited.

Sources

docsMDN Promise.all concurrency and rejection behaviordeveloper.mozilla.orgdocsMDN HTTP cachingdeveloper.mozilla.orgdocsPostgreSQL EXPLAIN and execution analysispostgresql.orgdocsPostHog engineering work-sample guidanceposthog.com

Checkpoint

Session 100 precedes independent 250/300 calls; render 80 follows both. Estimate?

A480 ms.B380 ms.C650 ms.D730 ms.
Sign up free to answer and see why

Checkpoint

Activity requires project version. What follows?

AIgnore activity failures to claim the faster time.BParallel start is valid because both use one project ID.CStart activity with any cached version without a policy.DThe dependency must remain or be replaced by a defined consistent contract.
Sign up free to answer and see why

Checkpoint

Warm cache saves 130 ms while client stages remain 880 ms. Best next investigation?

AOnly extend cache lifetime.BOnly change the loading message.CMeasure and reduce unnecessary payload/transformation/render work while preserving result semantics.DOnly increase database replicas.
Sign up free to answer and see why

Checkpoint

What does Promise.all rejection guarantee about other started work?

AIt does not automatically cancel the other operations.BEvery other request has stopped remotely.CNo side effects occurred anywhere.DAll results remain available as fulfilled values.
Sign up free to answer and see why

Checkpoint

A global total is calculated from only the first paginated page. Main issue?

APagination necessarily slows every endpoint.BThe total's population changed, so speed came with a correctness defect.CRendering must include all rows to compute any total.DCaching always repairs the total.
Sign up free to answer and see why

Can you distinguish total work from critical-path delay and select a valid measured improvement without changing correctness? State the relevant identifiers, failure boundary, and evidence in your own words before selecting your confidence.

Not yetGetting thereConfident

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.

Explain the waterfall with numbers · Full-stack interview…