Lesson 2 of 4 · 35 min

Concurrency is a capacity decision

Calculate a safe initial worker limit and identify what measurement can change it.

A queue absorbs a burst by delaying work. It does not create processing capacity. If arrivals remain above completions, backlog grows until storage, waiting time, or business deadlines become unacceptable. Worker count therefore needs a workload model and a downstream limit, not a guess based on CPU availability.
Start with units. If a task spends an average two seconds holding one provider request, ten concurrent tasks can complete roughly five tasks per second while the provider is healthy. This is a planning estimate, not a latency guarantee. Variance, startup work, rate limits, and retries reduce usable capacity. The relationship between average work in progress, throughput, and time is most useful over a stable observation window. During rapid backlog growth, treat it as a diagnostic approximation.
A concurrency limit and a rate limit protect different things. Concurrency caps simultaneous work. A rate limit caps starts or requests over time. A fast downstream service can receive too many requests per second despite low concurrency. A slow service can exhaust open connections despite a low request rate. A production worker may need both limits.
Reserve capacity for recovery and interactive traffic. A shared provider ceiling includes other callers. If the provider permits twenty requests per second and another application uses eight, assigning all twenty to batch work guarantees conflict. Backlog age is often more useful than queue length. A thousand jobs that take ten milliseconds differ from a thousand jobs that take ten minutes.

Worked example

A fictional provider allows 12 requests per second. Other traffic needs 4. Each job makes two provider calls, so the batch budget of 8 calls per second permits at most 4 jobs per second before retries. Measured average job duration is 1.5 seconds. A starting concurrency of 6 supports approximately 4 jobs per second, while a separate limiter enforces the call budget. Use a lower initial target if measurement is uncertain.
A burst adds 600 jobs while normal arrivals continue at 2 jobs per second. At sustained completion of 4 jobs per second, spare capacity is 2. Clearing the additional backlog takes about 300 seconds. Dividing 600 by 4 would incorrectly ignore new work arriving during recovery.

Build a capacity worksheet

Use a worksheet to keep rates and units visible. The following values are invented for a provider-backed export service. They are inputs to a planning model, not provider benchmarks.
QuantityValueUnit
Provider ceiling12calls per second
Reserved interactive traffic4calls per second
Calls per export2calls per job
Average job duration1.5seconds
Batch completion ceiling before retries4jobs per second
Starting concurrency estimate6active jobs
The rate ceiling follows from eight available calls divided by two calls per job. The concurrency estimate follows from four jobs per second multiplied by 1.5 seconds per job. This average relationship is useful only when the system is stable enough for the observation window to represent normal behavior. It does not establish a percentile latency or guarantee how a burst will behave.
A concurrency cap of six alone does not enforce eight calls per second. If provider responses suddenly become faster, six workers can start more calls each second. Keep a separate rate limiter for the provider budget. Conversely, a rate limiter alone can allow too many simultaneous slow calls if each call remains open for a long time. The two limits protect different resource dimensions.

Compare healthy and degraded windows

code
1Window H:2  arrivals=2 jobs/s3  completions=4 jobs/s4  extra backlog=600 jobs5  net drain=2 jobs/s -> estimate 300 s67Window D:8  arrivals=2 jobs/s9  completions=2 jobs/s10  extra backlog=600 jobs11  net drain=0 jobs/s -> backlog does not shrink1213Window O:14  arrivals=3 jobs/s15  completions=2 jobs/s16  net growth=1 job/s -> 60 additional jobs per minute
These windows explain why a queue length by itself is incomplete. A stable backlog of six hundred jobs might be draining, stuck, or growing depending on the rates. Oldest pending age adds information about whether some work never progresses. A queue can have healthy average throughput while one poison job or one low-priority tenant remains stuck.
Measure service-time variation. If most jobs take one second but a small fraction take five minutes, a single average can hide occupied slots and user deadlines. Separate job classes where their resource needs differ. A small export and a large archive may need different concurrency pools or admission rules. Do not use arbitrary worker counts for every class and assume the combined load remains within the shared provider ceiling.

Reserve recovery capacity deliberately

A recovering provider can be overwhelmed by the stored backlog plus new arrivals plus retries. Decide whether to reduce new admission, slow retries, or temporarily prioritize old work. The correct policy depends on user deadlines and fairness. A system promising interactive completion may need to reject new work with a clear retry path rather than accept jobs it cannot finish within the contract.
The worksheet should include retry calls. If twenty percent of jobs need one additional call, average calls per completed job become 2.2 under that simplified model. Eight available calls per second then support about 3.64 jobs per second, before other constraints. Six active workers may still be appropriate, but the rate limiter becomes the binding ceiling.

Misconceptions to correct

The first misconception is that doubling workers doubles throughput. It can increase throughput only while the limiting resources have spare capacity. Once the provider rate is binding, more workers can increase waiting and errors without delivering more completed jobs.
The second misconception is that queueing makes overload harmless. A queue trades immediate rejection for waiting and stored responsibility. If arrival rate remains above completion rate, the delay grows without bound unless admission or capacity changes.

Extend the exercise

Use a provider budget of nine calls per second, three calls per job, and continuing arrivals of two jobs per second. The maximum modeled completion rate is three jobs per second. A burst of 240 jobs drains at one job per second, so the estimate is 240 seconds. Award one point for each rate and one for distinguishing the estimate from observed recovery. Then ask which measurement would show that the provider is no longer the limiting resource.

Exercise and solution

The same provider slows, making jobs take 3 seconds. Keep concurrency at 6 and assume the call budget remains sufficient. Throughput falls to about 2 jobs per second. With normal arrivals at 2, the burst stops shrinking. The model answer does not immediately double concurrency. It first verifies provider saturation and the cause of slower service. Award two points for the arithmetic and one for naming the downstream constraint.

Interview probe and wrap-up

What metric would justify raising concurrency? A strong answer combines spare downstream capacity, latency behavior, error rates, and backlog age. Follow up with tenant-level skew. A weak answer scales workers until the queue disappears. Capacity decisions must account for the full path, including the service that the worker does not control.

Sources

docsAWS exponential backoff and jitteraws.amazon.comdocsMDN HTTP 429 statusdeveloper.mozilla.orgdocsAWS avoiding insurmountable queue backlogsd1.awsstatic.com

Checkpoint

Why is a concurrency cap insufficient as a provider rate limit?

AConcurrency always implies one call per second.BRate and concurrency use the same units.COnly queue length controls call rate.DFaster calls can increase starts per second at the same concurrency.
Sign up free to answer and see why

Checkpoint

Nine calls/s are available and each job uses three calls. Ceiling before retries?

ANine jobs/s.BThree jobs/s.CSix jobs/s.DOne job/s.
Sign up free to answer and see why

Checkpoint

Completion is three jobs/s and arrivals are two. How long to drain 240 extra jobs?

A240 seconds.B480 seconds.C80 seconds.D120 seconds.
Sign up free to answer and see why

Checkpoint

When completion equals arrival rate, what happens to an existing extra backlog?

AIt disappears after one lease period.BIt necessarily doubles.CIt does not shrink under the stated rates.DIt clears at the completion rate.
Sign up free to answer and see why

Checkpoint

Why inspect oldest pending age alongside queue count?

AIt proves that every queued job will meet its deadline.BIt can reveal starvation or stuck work despite healthy averages.CIt replaces all throughput measurements.DIt proves every job has equal cost.
Sign up free to answer and see why

Can you derive separate rate and concurrency limits, then explain what happens to backlog under healthy, flat, and overloaded windows? State the relevant identifiers, failure boundary, and evidence in your own words before selecting your confidence.

Not yetGetting thereConfident

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.