Revise a small architecture using measured workload rather than scale slogans.
A design discussion becomes more revealing when the interviewer changes a constraint. The goal is not to defend the original diagram at any cost. Identify which assumption changed, which guarantee is now threatened, and the smallest change that restores the required behavior. Keep unrelated parts of the design stable unless evidence connects them.
Growth can mean more users, more work per user, more variation, or a stricter latency target. These are different pressures. A thousand users who export once a month can create less load than ten users running continuous large exports. Describe rates, sizes, concurrency, and deadlines rather than using customer count as a direct infrastructure requirement.
A measured bottleneck points toward a specific change. If workers saturate a provider limit, more application servers may not help. If interactive requests wait behind CPU-heavy jobs, separating worker processes and capping concurrency may be enough. If one database write invariant is highly contended, partitioning unrelated work can help while the hot resource still needs a correct allocation strategy.
Reliability requirements also change cost. A pilot that tolerates manual recovery differs from a contract requiring rapid unattended recovery. That may justify more automation, redundancy, and operational coverage. State who responds and what happens outside working hours. A diagram with two regions does not by itself provide an operated service.
Worked example
A fictional reporting product grows from 5 to 60 exports per minute. Each export holds a worker for 10 seconds. At the new rate, average active work is about 10 jobs under stable conditions. The current worker cap is 4, giving approximate capacity 24 exports per minute. Backlog grows by 36 per minute before other limits.
The first proposal is to measure provider capacity and worker resource use, then increase independent worker capacity if the downstream path supports it. The web application remains one deployment because its measured latency is healthy. If the provider ceiling supports only 30 exports per minute, the product needs admission limits, scheduling, or a different service contract; more workers cannot create missing provider capacity.
Make the workload model inspectable
The example's average concurrency estimate follows from throughput multiplied by time in the system under stable conditions. Sixty exports per minute is one per second. At ten seconds of occupied worker time each, approximately ten workers are busy on average if the system keeps up. That is a baseline, not a sufficient production capacity plan. Variation, retries, scheduling, and a utilization margin can require more capacity.
Quantity
Current cap
New demand
Worker slots
4
At least 10 average occupied slots under supplied model
Service time
10 seconds
10 seconds assumption
Completion capacity
4×60/10=24 per minute
Demand 60 per minute
Net backlog growth
60-24=36 per minute
Before retries or provider limits
Write the assumptions with the numbers. If service time rises when concurrency increases, linear extrapolation fails. If workers spend most time waiting on a provider, CPU utilization may remain low despite capacity being exhausted. Measure the actual blocking resource.
A backlog of 360 jobs at net growth 36 per minute forms in ten minutes under the simple constant-rate model. That does not imply every user waits ten minutes, because queue discipline and job sizes matter. Oldest-job age and per-class completion provide additional evidence.
Introduce a downstream ceiling
Suppose the provider permits thirty exports per minute. Adding twenty worker slots cannot produce sixty legitimate completions per minute through that provider contract. It may increase throttling, contention, or retries. The product needs to reduce admitted rate, schedule work, negotiate a supported service level, change workload, or select an alternative that meets requirements.
code
1Changed constraint record2Demand:60 exports/minute3Worker capacity after proposed scale:120/minute4Provider ceiling:30/minute5Effective sustained path capacity:at most30/minute6Required decision:7 cap admission, extend deadline, reduce provider work,8 or change the supported dependency contract
This record separates a mechanism from a business decision. If the customer was promised immediate completion, extending the queue delay changes the promise and needs explicit agreement. If the task is a nightly batch due by morning, scheduling may meet the outcome with lower cost. The architecture cannot choose the deadline silently.
Change reliability instead of traffic
Now keep demand at five exports per minute but require recovery without the sole operator being awake. The pressure is operational, not throughput. Durable job state, bounded retries, automatic restart of a failed worker, alerting, and a tested recovery path may matter more than splitting the application into services.
Define which failures the product must recover from automatically and how quickly. A process restart does not repair an invalid input or an unknown provider effect. Automation needs failure classification and stable operation identity. A design that retries every failure indefinitely is neither reliable nor bounded.
A standby operator or provider support agreement can be part of the solution when the promise requires human decisions. Do not draw a redundant component and imply that staffing, runbooks, and data compatibility are solved. Reliability is a delivered operating capability.
Explain the minimal sufficient change
In a design interview, say which parts stay valid because their assumptions did not change. If the web path is healthy, more web replicas are not the first response to a provider-limited export queue. If the failure is cross-account cache reuse, worker scaling does not address the correctness boundary. Keeping unrelated parts stable is a reasoned choice, not resistance to growth.
A migration plan also needs a verification sequence. Introduce the change for a small representative workload, compare completion and oldest age, monitor provider outcomes, and preserve a rollback or stop boundary. Scaling concurrency without observing downstream effects can amplify the original problem.
Misconceptions and a second exercise
One misconception is that customer count directly determines service count. Work per customer and concentration matter. Another is that average capacity exactly equal to average arrivals guarantees low latency. Bursts and variability can grow queues even when long-run averages appear balanced.
Exercise: demand averages twenty jobs per minute, but all arrive in a burst at the start of each minute. Each takes ten seconds and four slots are available. Five waves are needed, so the last wave finishes around fifty seconds after arrival under equal-duration/no-overhead assumptions. Average capacity is twenty-four per minute, yet the last users still wait. Award one point for five waves, one for fifty seconds, one for identifying the burst effect, and one for relating the delay to the promised deadline.
The best answer names the changed assumption, computes the relevant limit, proposes a bounded response, and states what measurement could disprove it. This is stronger than adding components until the diagram looks large.
Exercise and solution
The interviewer instead says export volume is unchanged, but private customer results are mixed in a shared cache. What should change first? Fix cache identity and authorization-related representation, and contain exposure. Scaling workers is irrelevant to that failure. Award one point for identifying correctness rather than capacity, one for a valid boundary fix, and one for a regression scenario across two accounts.
Interview probe and wrap-up
When would you accept slower service? A strong answer ties the choice to the user deadline, budget, and communicated queue behavior. Follow up with a paid promise of immediate results. A weak answer says the system must always be real time. A good design adapts to the changed constraint while making the cost and remaining limits explicit.
Demand 60/minute and provider ceiling 30/minute. More workers alone?
AGuarantee 60 completions/minute.BEliminate the need for queue policy.CProve the database is the bottleneck.DCannot remove the provider ceiling and may increase throttling.
Traffic is unchanged but unattended recovery becomes required. What should be reviewed first?
AFailure classification, durable state, bounded automation, and operational coverage.BOnly increase worker slots because recovery is the same as capacity.CKeep manual recovery but describe retries as automatic recovery.DRetry every failure indefinitely without classification.
Can you adapt architecture to rate, burst, dependency, or reliability changes without confusing these different constraints? State the relevant identifiers, failure boundary, and evidence in your own words before selecting your confidence.