Lesson 2 of 4 · 25 min

Budget requests, limits and scheduling

Calculate schedulable resource demand and diagnose enforcement.

Mechanism and reasoning

A resource request helps the scheduler decide where a Pod can fit. A limit controls runtime consumption through the relevant enforcement mechanism. These are different jobs. A low request can make a workload easy to schedule while leaving it vulnerable to contention. A high request can keep it pending even when observed use appears low.
CPU is compressible in the sense that a workload can be throttled rather than immediately terminated when it reaches an enforced CPU limit. Memory is different. A process that exceeds available or enforced memory can be killed. The exact outcome depends on runtime and node conditions, so inspect termination reasons and resource metrics instead of guessing from a restart count.
The arithmetic in this lesson assumes ordinary, concurrently running containers with container-level requests. It excludes init-container scheduling rules, Pod-level resource budgets, and additional Pod overhead unless a case supplies them. Inspect those rules before applying a simple sum to another Pod shape. Scheduling uses allocatable resources, not the node's advertised total. The operating system and cluster components need a reserve. Include every container in the Pod and the applicable overhead. Sidecars can materially change the request. A table that counts only the main application understates demand.
Requests also influence utilization-based autoscaling. If CPU use is compared with requests, changing requests changes the reported percentage even when actual CPU stays constant. A configuration edit can therefore alter both placement and scaling behavior. Explain these interactions before adjusting values to make a dashboard look better.
Choose resource values from observed workload distributions and the service objective. A peak memory request for every Pod may be expensive; a tiny request with a large limit may create contention risk. There is no universal correct ratio. Measure startup, normal load and the worst supported operation, then document the chosen reserve and failure behavior.

Placement worksheet

Teaching cluster: two nodes each expose four allocatable CPU cores and eight GiB of allocatable memory. Each application Pod has a main container requesting one core and two GiB, plus a sidecar requesting 0.25 core and 0.5 GiB. The Pod therefore requests 1.25 cores and 2.5 GiB. One node can host at most floor(4/1.25)=3 by CPU and floor(8/2.5)=3 by memory. The simplified capacity is three Pods per node, or six across the cluster.
Now add a constraint that all Pods require a particular node label, and only one node has it. The arithmetic still says six in aggregate, but only three are eligible. Placement constraints, taints, volumes and topology can reduce usable capacity. Resource sums are necessary evidence, not a complete scheduler simulation.
Suppose the service needs four ready Pods during a node failure. Two nodes with three each cannot meet that requirement because one loss leaves only three. Add a third eligible node or change the workload shape if measurements justify it. Reducing requests merely to make the arithmetic pass is not a capacity improvement. It changes the scheduler's promise while actual work can remain unchanged.
The next table separates several similar-looking symptoms. A Pending Pod has not reached normal execution and may lack an eligible placement. A running container with high CPU throttling can have a limit problem or an intentional cap. An OOMKilled container points toward memory enforcement or pressure and needs its own evidence. The same response, 'increase replicas,' does not fit all three.
SymptomRelevant evidenceCandidate correction
Pending with insufficient CPUPod requests, eligible nodes, allocatable CPUAdd eligible capacity or revise justified requests
Running with throttling and latencyCPU use, throttling, limits, loadReassess limit and workload capacity
OOMKilled during importPeak memory, limit, import sizeBound the operation or provide justified memory
Low average use but failed placementRequests and constraintsExplain reserved capacity instead of using averages

Resource contract and failure test

A resource contract should state the supported workload. For example, an API may support imports up to fifty thousand rows with a measured memory peak below 1.5 GiB. An unlimited import endpoint cannot be made safe merely by assigning a two-GiB limit. The limit bounds damage to the node but can turn an oversized request into a restart. Add an application limit or a streaming algorithm and test the largest allowed input.
Test one Pod under representative load, then test several Pods sharing a node. The single-Pod result does not expose contention. Include startup because initialization can exceed steady-state memory. Keep the input and concurrency fixed when comparing request changes. Otherwise a lower resource bill can reflect less work rather than better efficiency.
Finally, distinguish observed utilization from reserved capacity in an interview answer. A node showing twenty percent CPU can still lack room for another requested Pod. The scheduler protects declared requests, not just the current instantaneous graph. This can look wasteful, but changing it requires a workload and risk argument rather than a screenshot.

Worked example

Teaching Pod requests 750 millicores for the app and 250 for a sidecar, totaling one core. On a node with 3.5 allocatable cores and sufficient memory, three such Pods fit by CPU. A fourth requires four cores and remains unschedulable under this simple model. If the app uses 500 millicores and an autoscaler compares use with a 750-millicore request, that container's utilization is about 66.7%. Lowering the request to 500 changes the ratio to 100% without changing actual work.

Exercise

A node has six allocatable cores and twelve GiB. Each Pod requests 1.5 cores and two GiB, including all containers. Find the maximum by resources. Then state remaining capacity after placing three Pods and explain why memory alone cannot justify a fifth Pod.

Model solution and rubric

CPU permits four Pods and memory permits six, so CPU limits placement to four. After three Pods, 1.5 cores and six GiB remain. A fourth fits. A fifth would require 7.5 cores total, exceeding six despite enough memory for its ten-GiB total. Verify other scheduler constraints and node overhead assumptions before declaring actual fit. A lower observed CPU average does not override the request-based placement calculation.
Score out of four: one point for the correct result, one for showing the intermediate reasoning, one for identifying the stated failure case, and one for a verification that could disprove the answer. Do not award the reasoning point for a tool name alone.

Failure modes and misconceptions

Misconception 1: limits reserve capacity. Requests drive the relevant scheduling reservation; a high limit alone does not secure that capacity. Misconception 2: an idle-looking node must accept a Pending Pod. Requests, eligibility and topology can prevent placement even when instantaneous use is low.

Interview probe

Evidence class: recommended. Original practice.
Why is a Pod Pending when the node is only 20% busy?
Strong answer: I inspect requested resources and placement constraints against allocatable capacity. The scheduler does not place from the current CPU graph alone. I also include sidecars, taints, labels and volume constraints.
Follow-up: How can lowering CPU requests unexpectedly change autoscaling?
Weak answer indicators: Using advertised capacity instead of allocatable; ignoring sidecars; treating CPU throttling and memory kills as the same mechanism.

Sources

Technical references: Kubernetes container resources; Kubernetes Deployments. Sources support the documented mechanisms. The numbers, decisions, rubrics and interview prompts in this lesson are original teaching examples, not measurements or employer question claims.
docsKubernetes container resourceskubernetes.iodocsKubernetes Deploymentskubernetes.io

Checkpoint

A node has 2 CPU allocatable, 1.8 CPU already requested, and only 0.4 CPU in current use. A new Pod requests 0.5 CPU. With no other eligible node, what does the scheduler use for this fit check?

AThe 0.4 CPU current use, so the new Pod fitsBThe current request total, so 1.8 + 0.5 exceeds 2 CPUCThe sum of CPU limits, regardless of requestsDThe average of requests and measured use
Sign up free to answer and see why

Checkpoint

Three eligible nodes each have eight allocatable CPUs. Every application Pod requests three CPUs. One node fails. With no other workloads or placement limits, can the two surviving nodes hold five application Pods?

ANo. Each node fits two Pods, so only four fit despite sixteen total CPUs.BYes. Five Pods need fifteen CPUs, less than the surviving aggregate of sixteen.CYes. The failed node's reservations remain usable until its Pods restart.DNo. Every surviving node can hold only one Pod because a third would exceed eight CPUs.
Sign up free to answer and see why

Checkpoint

At fixed request rate, a container reaches its 0.5 CPU limit and shows throttling while node CPU has spare capacity. p99 latency rises; memory and database timing remain stable. Which next comparison best tests the suspected cause?

AIncrease memory and compare under a lower request rateBIncrease the CPU request only while leaving the 0.5 CPU limit unchangedCUse a controlled higher CPU limit at the same load and compare throttling plus end-to-end latencyDRestart the container and compare only its first request
Sign up free to answer and see why

Checkpoint

Actual CPU stays 400m while request drops from 800m to 400m. Utilization relative to request becomes?

A25%B50%C75%D100%
Sign up free to answer and see why

Checkpoint

A main container fits but its sidecar pushes total requests over capacity. What should the plan use?

AMain container onlyBSum of applicable Pod requests and overheadCThe largest limit onlyDThe average request across containers
Sign up free to answer and see why

Explain how you would calculate schedulable resource demand and diagnose enforcement without reading the solution. State one assumption that could change your answer, and one observation that would make you revise it.

Not yetGetting thereConfident

Wrap-up

  • Separate placement from enforcement. Carry every container and failure requirement through the capacity calculation.

Sources

Free to read · better with Enzo

Learn it with Enzo

Save your progress, answer the checkpoints, and let Enzo quiz you on what you just read.