Kubernetes CPU and Memory Requests vs Limits
Kubernetes CPU and Memory Requests vs Limits
bitcodematrix.com | Kubernetes Series
1. Introduction
Every container in a Kubernetes cluster competes for the same finite CPU and memory on a node. Resource requests and limits are how you tell Kubernetes how much a container needs and how much it is allowed to use. They look like two numbers in a YAML file, but they drive scheduling, performance, stability, autoscaling and even which Pods are killed first when a node runs out of memory.
Many production incidents trace back to getting these values wrong: Pods stuck in Pending, containers killed with OOMKilled, services that slow down for no obvious reason, or nodes that are oversubscribed and unstable. This article explains what requests and limits really do, how CPU and memory behave differently, how Quality of Service classes work, and how to choose sensible values.
After reading it you should be able to:
Explain the difference between a request and a limit.
Describe how CPU is throttled while memory is enforced by the OOM killer.
Read and write CPU and memory quantities such as 250m and 512Mi correctly.
Explain the three QoS classes and the eviction order.
Choose, measure and tune values, and troubleshoot related failures.
2. Requests and Limits at a Glance
A request is the amount of a resource that Kubernetes reserves for a container. A limit is the maximum the container is allowed to use. They are set per container, in the resources section of the container spec.
Figure 1: The scheduler reads requests; the kubelet and cgroups enforce limits
3. Units: How to Write CPU and Memory
3.1 CPU
CPU is measured in CPU units. One unit equals one physical core or one virtual core, depending on the node. You can use decimals or the millicore suffix m, where 1000m equals 1 CPU. The smallest precision is 1m.
3.2 Memory
Memory is measured in bytes, usually written with a suffix. Be careful with the difference between binary and decimal suffixes.
As a rule, use Mi and Gi for memory. A very common typo is writing 512m for memory when 512Mi was meant, which results in a container that cannot start or is killed immediately.
4. How Requests Work
4.1 Scheduling
When a Pod is created, the scheduler adds up the requests of all its containers (and init containers, using the larger of the sum of regular containers and the biggest init container, as per the documented rules) and looks for a node whose unreserved allocatable capacity can fit them. The scheduler looks only at requests. It does not care how much a Pod actually uses at that moment, so a node can be nearly idle and still reject new Pods because all of its capacity is already requested.
Allocatable capacity is not the same as total node capacity. It is the capacity left after reserving resources for the operating system, the kubelet and the container runtime, plus eviction thresholds.
4.2 CPU requests at runtime
On Linux, the kubelet translates the CPU request into a cgroup weight (cpu.shares on cgroup v1, cpu.weight on cgroup v2). This weight only matters when the CPU is contended. If a node has spare CPU, any container can use it. If several containers want more CPU than is available, each receives CPU in proportion to its request. A container requesting 500m gets roughly twice the share of one requesting 250m when both are competing.
4.3 Memory requests at runtime
Memory requests do not reserve physical memory for the container in a hard sense. They are used by the scheduler for placement and by the kubelet when deciding which Pods to evict under node memory pressure. A Pod using more memory than it requested is a more likely eviction candidate than one staying below its request.
5. How Limits Work
5.1 CPU limits: throttling
CPU is a compressible resource: if a container wants more than its limit, Kubernetes does not kill it; it slows it down. The limit is enforced with the Linux Completely Fair Scheduler (CFS) quota mechanism. Time is divided into periods (100 ms by default), and the container receives a quota of CPU time in each period. A limit of 500m means 50 ms of CPU time per 100 ms period (spread across all of the container's threads). Once the quota is used up, the container's threads are paused until the next period.
This means a multi-threaded application with a low limit can be throttled even when its average CPU usage looks well below the limit, because bursts consume the quota quickly. Throttling shows up as increased latency, not as errors or restarts, which makes it hard to notice without the right metric.
5.2 Memory limits: OOM kill
Memory is an incompressible resource: it cannot be taken back without ending the process that holds it. If a container tries to use more memory than its limit, the kernel's out-of-memory (OOM) killer terminates a process in the container. When the main process is killed, the container is reported as OOMKilled with exit code 137, and the kubelet restarts it according to the Pod's restartPolicy. Repeated OOM kills lead to CrashLoopBackOff.
Note that this is a container-level OOM kill triggered by the cgroup limit. A node-level shortage is handled differently, by kubelet eviction and the kernel OOM killer acting on the whole node, as explained in section 7.
5.3 Summary of behaviour
6. Writing Requests and Limits
apiVersion: v1
kind: Pod
metadata:
name: web
spec:
containers:
- name: app
image: nginx:1.27
resources:
requests:
cpu: 250m
memory: 256Mi
limits:
cpu: 500m
memory: 512Mi
In this example the scheduler needs a node with at least 250m CPU and 256Mi memory unreserved. At runtime the container may burst up to half a core, and it will be killed if it goes beyond 512Mi of memory.
Pod-level totals are the sum across containers. A Pod with two containers each requesting 250m CPU has a Pod request of 500m. You can inspect what a node has already committed with the following command:
kubectl describe node <node-name>
# Look for the section:
# Allocated resources:
# Resource Requests Limits
# cpu 1850m (46%) 4200m (105%)
# memory 2100Mi (27%) 5000Mi (64%)
The values above are illustrative. Note that the limits total can exceed 100% of node capacity. This is called overcommitment and is allowed for limits, but not for requests, because the scheduler will not place a Pod if the requests would exceed allocatable capacity.
7. Quality of Service (QoS) Classes
Kubernetes assigns every Pod one of three QoS classes, based on the requests and limits of its containers. You cannot set the class directly; it is derived.
You can see the class of a running Pod:
kubectl get pod web -o jsonpath='{.status.qosClass}'
7.1 Why QoS matters: eviction and OOM order
When a node runs low on memory, the kubelet evicts Pods to reclaim it. If memory runs out faster than the kubelet can react, the kernel OOM killer steps in. In both cases the QoS class and resource usage relative to requests influence the order:
BestEffort Pods are the first candidates.
Burstable Pods that use more than their memory request come next, with the largest overage generally ranked higher for eviction (Pod priority is also considered by the kubelet).
Guaranteed Pods, and Burstable Pods using less than their requests, are protected the longest.
The kernel also receives an oomscoreadjust value from the kubelet, strongly favouring Guaranteed Pods for survival and BestEffort Pods for termination. Even Guaranteed Pods can still be evicted if system daemons are starved or the node is truly out of resources.
7.2 CPU Manager and Guaranteed Pods
If the kubelet is configured with the static CPU Manager policy, Guaranteed Pods that request a whole number of CPUs (for example cpu: "2" with matching limit) can be given exclusive cores. This is useful for latency-sensitive or NUMA-aware workloads, but it requires deliberate node configuration.
8. Choosing Good Values
8.1 Measure first
Do not guess. Run the application under realistic load and observe actual usage over days, including peaks. Useful sources are kubectl top (which needs metrics-server), Prometheus metrics, and the Vertical Pod Autoscaler in recommendation mode.
kubectl top pod -n shop
kubectl top pod web --containers
8.2 Practical sizing guidance
8.3 The CPU limit debate
Teams disagree about whether to set CPU limits at all, and both positions are reasonable.
Argument for limits: they cap noisy neighbours, make performance predictable across environments, are required for the Guaranteed QoS class, and may be enforced by quotas or policy.
Argument against limits: because spare CPU is otherwise wasted, limits can cause needless throttling and latency on a node with idle capacity. CPU requests alone already guarantee a fair share when the node is busy.
A balanced approach is to always set CPU requests, set CPU limits where your platform policy or QoS requirements demand them, make those limits generous, and monitor throttling. Memory is different: always set a memory limit, because without one a single leak can destabilise the whole node.
8.4 Runtime-specific tips
JVM: modern JVMs read the container memory limit. Use options such as -XX:MaxRAMPercentage so that heap plus non-heap memory stays below the limit; a heap set equal to the limit will be OOM-killed.
Node.js, Python, Go: leave headroom above the language heap for buffers, native memory and runtime overhead. For Go, GOMEMLIMIT can help the garbage collector respect the container limit.
Thread pools: size worker counts relative to the CPU limit, not the node's core count, otherwise heavy throttling can result.
9. Interaction with Other Features
Resizing requests and limits of a running Pod without restarting it is available through in-place Pod resize in recent Kubernetes versions. Its maturity level depends on the version, so check the documentation for the release you run before relying on it.
The article on ResourceQuotas and LimitRanges in this series covers namespace-level controls in more detail, and the article on security best practices explains why limits are also a safety control.
10. Hands-On Walkthrough
This exercise demonstrates a Pending Pod, an OOM kill and CPU throttling. Use a test cluster with metrics-server installed.
10.1 Trigger a Pending Pod
# Apply a manifest asking for an unrealistic request
cat <<EOF | kubectl apply -f -
apiVersion: v1
kind: Pod
metadata:
name: too-big
spec:
containers:
- name: app
image: nginx:1.27
resources:
requests:
cpu: "100"
memory: 1Ti
EOF
kubectl get pod too-big
kubectl describe pod too-big
The Pod stays in Pending, and the Events section shows a FailedScheduling message similar to "0/3 nodes are available: 3 Insufficient cpu, 3 Insufficient memory". The exact text varies by version and cluster.
10.2 Trigger an OOM kill
cat <<EOF | kubectl apply -f -
apiVersion: v1
kind: Pod
metadata:
name: oom-demo
spec:
containers:
- name: stress
image: polinux/stress
command: ["stress"]
args: ["--vm", "1", "--vm-bytes", "300M", "--vm-hang", "1"]
resources:
requests:
memory: 100Mi
limits:
memory: 200Mi
EOF
kubectl get pod oom-demo -w
kubectl describe pod oom-demo | grep -A5 "Last State"
Expected result: the container is terminated with Reason OOMKilled and Exit Code 137, restarts, and eventually shows CrashLoopBackOff. The image is a commonly used community stress tool; substitute any equivalent image of your choice.
10.3 Observe CPU throttling
Run a CPU-heavy container with a low CPU limit and compare usage and throttling metrics.
cat <<EOF | kubectl apply -f -
apiVersion: v1
kind: Pod
metadata:
name: cpu-demo
spec:
containers:
- name: stress
image: polinux/stress
command: ["stress"]
args: ["--cpu", "2"]
resources:
requests:
cpu: 100m
limits:
cpu: 200m
EOF
kubectl top pod cpu-demo
Two busy threads would normally use two cores, but kubectl top shows usage held near 200m. In Prometheus, compare the following metrics (names as exposed by cAdvisor):
rate(containercpucfsthrottledperiodstotal[5m])
/ rate(containercpucfsperiods_total[5m])
A high ratio means the container is frequently being throttled. Clean up afterwards with kubectl delete pod too-big oom-demo cpu-demo.
11. Troubleshooting
12. Best Practices
Always set CPU and memory requests for every production container.
Always set a memory limit; consider setting memory request equal to limit for predictability.
Treat CPU limits as a deliberate decision: if you set them, make them generous and monitor throttling.
Base values on measured usage, and revisit them as the application changes.
Use Mi and Gi for memory and double-check suffixes.
Use LimitRanges to provide defaults and ResourceQuotas to bound each namespace.
Use Guaranteed QoS for the most critical workloads, and avoid BestEffort in production.
Leave headroom above the language runtime heap so the container does not hit its memory limit.
Monitor OOM kills, evictions, throttling, and Pending Pods, and alert on them.
Consider the Vertical Pod Autoscaler in recommendation mode to keep values realistic.
13. Conclusion
Requests and limits answer two different questions. A request says "how much should you set aside for me?" and is used for scheduling, fair sharing and eviction ranking. A limit says "how much may I use at most?" and is enforced by throttling for CPU and by the OOM killer for memory. Because CPU is compressible and memory is not, the two resources deserve different treatment: be careful with CPU limits, and always protect the node with memory limits. Start from measurements, choose the QoS class that fits each workload's importance, and keep watching the metrics so that your numbers stay true as your applications evolve.