Lesson 2 of 4 · 35 min

Equal steps may not mean a fair resource comparison

Select and report a resource comparison that matches the claim.

Two methods can use the same number of training steps while processing different numbers of examples, tokens, or floating-point operations. They can use equal compute while taking different wall time. A fair comparison starts by stating the resource question, not by assuming one universal budget unit.
A fixed-step comparison asks what happens after the same update count under the specified batch and data policy. A fixed-token comparison controls exposure more directly for language tasks. A fixed-compute comparison asks how effectively methods use computation. A time-to-quality comparison asks how quickly the system reaches a target metric on a specified environment. Each can be valid, but they support different claims.
Include tuning and failed runs when discussing total research cost. A final training run may be cheap only because a large search found its settings. If the claim concerns deployment inference cost, report the request distribution, batch size, latency constraints, and hardware. Training efficiency does not imply serving efficiency.
Use a Pareto view when quality and cost trade off. A method is dominated if another is at least as good in quality and no worse in the relevant cost, with one strict improvement. A non-dominated method is not automatically the best choice; the application still needs a preference or constraint. Report the available tradeoff instead of hiding it in a composite score with arbitrary weights.

Worked example

An invented comparison has three methods.
MethodValidation scoreTraining GPU-hours
A8010
B8220
C7915
C is dominated by A under these two measures: A has higher quality and lower cost. A and B are not mutually dominated. B gains two points for ten additional GPU-hours. A budget of 12 hours favors A; a quality requirement above 81 may require B or a new method.
Now add tuning cost: A used five hours of search, B used 100, and C used none. The final-run table remains factually correct, but a claim about total research efficiency needs the search costs too. Keep those quantities separate rather than changing the denominator only for the preferred method.

Exercise and solution

Method D reaches 85% after 100 minutes; E reaches 85% after 80 minutes but peaks at lower final quality. What comparison does E win, and what remains unresolved?
E wins time to the stated 85% target under the measured conditions. It does not necessarily win final quality, another target, or cost on different hardware. Award one point each for the bounded metric, environment condition, final-quality distinction, and a proposal to report the full learning curve. Do not summarize every efficiency claim as "faster" without naming the target.

Lab artifact: change the resource axis explicitly

The original table uses final training cost. Add total study cost and serving latency to see why one frontier is not universal.
MethodQualityFinal training hoursSearch hoursServing p95
A8010540 ms
B822010090 ms
C7915030 ms
Under quality and final training hours, A dominates C. Under quality and serving latency, A does not dominate C because C is faster. Under a service limit of fifty milliseconds, B is infeasible despite its higher measured quality. The claim must name its axes and constraints. “Dominated” is shorthand for a defined comparison, not an intrinsic property of a method.
Total study cost is fifteen hours for A, one hundred twenty for B, and fifteen for C under this simplified record. If earlier pretrained artifacts are reused, decide whether the study-cost claim includes their production cost or only marginal project cost. Both perspectives can be useful, but changing the accounting boundary only for one method is misleading. Disclose shared and sunk costs separately when they affect interpretation.

A second failure case: matching steps but changing token work

Method P processes four sequences of length one thousand per update. Method Q processes four sequences of length two thousand. Over ten thousand updates, P processes forty million sequence tokens and Q eighty million, before accounting for padding or repeated views. Equal updates do not imply equal input exposure. Even equal valid tokens may not imply equal compute because attention, architecture, and kernels can scale differently.
code
1Resource contract for a controlled comparison:2count unit: valid input tokens, excluding padding3budget: 40 million valid tokens4hardware: fixed device type and count5timing: includes data transfer and evaluation as specified6quality target: selected before inspecting learning curves7search budget: same number of permitted candidate evaluations8failed attempts: recorded and included in total study cost
This is one defensible contract, not the only fair design. A fixed-wall-time study may deliberately allow different token counts because the question is which recipe uses a time budget better. In that case, report both elapsed time and achieved exposure rather than calling the methods identical in work.

Exercise: apply a new deployment constraint

Methods U and V have quality scores eighty-three and eighty-four. U costs two units per request and takes sixty milliseconds p95; V costs three units and takes forty milliseconds. The application allows at most fifty milliseconds and four cost units. Which method is feasible under the stated measurements?
V passes both limits; U fails latency. This does not prove V is scientifically superior in every respect. If the latency limit changes to seventy milliseconds and cost becomes at most 2.5 units, U is feasible and V fails cost. Award one point for each decision, one for keeping conditions fixed, and two for distinguishing feasibility from universal superiority. Report uncertainty and measurement conditions if these are estimates near a boundary.

Misconceptions to correct

“Non-dominated means best” fails because choosing among tradeoffs requires a preference or constraint. “Same GPU-hours means same wall time” fails when devices run in parallel, differ in type, or spend time on host and communication overhead. A hardware-hour count is an accounting quantity with its own boundary.

Transfer calculation: total cost depends on deployment volume

A larger one-time study cost can buy a smaller serving cost. Under a declared linear model, let F include search plus final training, s be the measured cost per served request, and N the planned request volume. Then total cost is F + N×s. If A has larger F but smaller s than B, their break-even volume is (F_A-F_B)/(s_B-s_A). Compare methods only after checking their required quality, latency and safety constraints.
For an original example, A has F=500 and s=0.001; B has F=100 and s=0.003 in the same accounting units. They tie at 400/0.002=200,000 requests, when each costs 700 units. B is cheaper below this volume and A above it under the stated model. This does not imply a universal winner: uncertain demand, hardware reservations, idle capacity or changing batch efficiency can break linear per-request accounting. Include those costs or state their exclusion before using the result. Equal final training cost alone does not answer this deployment-lifetime question.
The resource report should include enough workload detail for another researcher to reconstruct the comparison: shape distribution, precision, batch size, target quality definition, timing policy, and tuning allowance. A profiler helps measure local behavior but does not itself define Pareto dominance or fairness. Those conclusions come from the declared objectives and the original comparison arithmetic.
A useful interview answer offers two views when the research question warrants them: quality at fixed resource and resource needed for fixed quality. If the methods' curves cross, show the crossing rather than selecting the favorable endpoint. The result can reveal that one method is preferable for short runs and another for higher final quality.

Interview probe

Original practice: Your method wins at equal steps but loses at equal compute. Which result should you report? A strong answer reports both and ties each to a different claim. Follow up with unequal tuning budgets. A weak answer chooses the favorable budget after seeing results.

Sources

docsNeurIPS: paper checklistneurips.ccdocsPyTorch: profiler guidedocs.pytorch.org

Checkpoint

A has quality 80 at training cost 10; C has 79 at cost 15. On these axes, what is true?

AA dominates C.BC dominates A.CThey are necessarily tied.DThe result proves A has lower serving latency.
Sign up free to answer and see why

Checkpoint

C serves in 30 ms while A serves in 40 ms. Does training-cost dominance establish serving-latency dominance?

AYes, every cost axis is interchangeable.BYes, because accuracy is higher.CNo; the relevant frontier depends on the named resource axis.DNo method can ever be dominated.
Sign up free to answer and see why

Checkpoint

P uses four 1,000-token sequences/update; Q uses four 2,000-token sequences/update. Over 10,000 updates, input-token exposure is?

ABoth 40 million.BP 4 million, Q 8 million.CBoth 80 million.DP 40 million, Q 80 million, before padding/view qualifications.
Sign up free to answer and see why

Checkpoint

A final run costs 20 hours after 100 search hours. What is the simplified total study cost?

A20 hours.B120 hours, with the accounting boundary stated.C100 hours.D80 hours.
Sign up free to answer and see why

Checkpoint

Under a comparison that maximizes quality and minimizes serving cost, what does being non-dominated establish?

AThe option is the unique preferred deployment choice for every quality/cost preference.BThe option satisfies any external latency constraint not included in the comparison.CNo compared option is at least as good on every named objective and strictly better on at least one; choosing among frontier options still needs preferences or constraints.DThe option must have the lowest cost among all compared options.
Sign up free to answer and see why

Can you choose a resource axis and apply feasibility or dominance without hiding search cost? Rate confidence from 1 to 5 and identify the accounting boundary.

Not yetGetting thereConfident

Wrap-up

  • State the resource axis and selection cost. A fair comparison answers the declared question rather than favoring one method.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.