Kubernetes Resource Troubleshooting: CPU, Memory and OOM Issues
Kubernetes Resource Troubleshooting: CPU, Memory and OOM Issues
bitcodematrix.com | Kubernetes Series
1. Introduction
Resource problems are among the most common and most confusing issues in Kubernetes. A Pod stays Pending although the cluster looks idle. A container restarts every few minutes with no error in its logs. A service becomes slow while its CPU graph looks calm. A node starts evicting Pods at night. The symptoms differ, but the causes all come back to the same few things: requests, limits, node capacity and how the Linux kernel enforces them.
This article is a practical troubleshooting playbook. It gives you a triage method, a toolbox of commands, and a walk through the four main symptom families: Pending Pods, CPU problems, memory problems and OOM kills, and evictions. It finishes with short case studies, a hands-on lab, a fix catalogue and a prevention checklist. It is the troubleshooting companion to the earlier articles on requests and limits, QoS classes, OOMKilled, scheduling and autoscaling.
After reading it you should be able to:
Classify a resource problem from its symptom within a minute or two.
Use kubectl, metrics and cgroup files to find the root cause.
Tell CPU throttling, node CPU starvation and application slowness apart.
Distinguish a container OOM kill, a node-level OOM and an eviction.
Apply the right short-term and long-term fix, and prevent recurrence.
2. A Quick Mental Model
Most resource troubleshooting becomes easier with five facts in mind.
3. A Triage Method
Figure 1: Triage by symptom
Ask these questions in order:
What is the Pod's state? Pending, Running with restarts, Failed or Evicted, or Running but slow.
What do the events and last state say? kubectl describe pod answers most questions: scheduling messages, Last State, Reason, Exit Code.
Is it one container or the whole node? Check whether other Pods on the same node are affected, and whether the node reports pressure conditions.
What do the numbers say? Compare requests, limits and actual usage (kubectl top and metrics).
What changed? New release, new traffic pattern, new limits, node upgrade, new neighbour workloads.
4. The Troubleshooting Toolbox
4.1 Kubernetes level
# State, restarts, node
kubectl get pod <pod> -n <ns> -o wide
# Events, Last State, Reason, Exit Code, requests and limits
kubectl describe pod <pod> -n <ns>
# Previous container instance logs (after a crash)
kubectl logs <pod> -n <ns> --previous
# Current usage (needs metrics-server)
kubectl top pod <pod> -n <ns> --containers
kubectl top node
# What the node has committed and its conditions
kubectl describe node <node>
# Recent events, newest last
kubectl get events -n <ns> --sort-by=.lastTimestamp
# QoS class and last termination reason
kubectl get pod <pod> -n <ns> -o jsonpath='{.status.qosClass}{"\n"}{.status.containerStatuses[*].lastState.terminated.reason}{"\n"}'
4.2 Inside the container (cgroup files)
Because the kernel enforces limits, the cgroup files show exactly what the container sees. The paths below are for cgroup v2, which most modern distributions use. On cgroup v1, the equivalents live under /sys/fs/cgroup/cpu and /sys/fs/cgroup/memory with different file names, such as memory.limitinbytes and memory.usageinbytes. The container image must contain a shell and cat for kubectl exec to work.
# CPU limit as "quota period" in microseconds, e.g. 50000 100000 = 0.5 CPU
kubectl exec <pod> -n <ns> -- cat /sys/fs/cgroup/cpu.max
# CPU usage and throttling counters
kubectl exec <pod> -n <ns> -- cat /sys/fs/cgroup/cpu.stat
# usageusec, nrperiods, nrthrottled, throttledusec
# Memory limit, current usage and OOM counters
kubectl exec <pod> -n <ns> -- cat /sys/fs/cgroup/memory.max
kubectl exec <pod> -n <ns> -- cat /sys/fs/cgroup/memory.current
kubectl exec <pod> -n <ns> -- cat /sys/fs/cgroup/memory.events
# oom, oomkill
If the image has no shell, use an ephemeral debug container that shares the target's process namespace, for example kubectl debug -it <pod> --image=busybox --target=<container>. Note that the cgroup files seen from a debug container may differ depending on the runtime, so treat them as a guide.
4.3 Node level and monitoring
# Kernel log entries for OOM kills (on the node)
journalctl -k | grep -i -E "out of memory|killed process"
# Prometheus examples (verify metric names in your setup)
# Working set against the limit
containermemoryworkingsetbytes{container!=""}
/ on(namespace,pod,container)
kubepodcontainerresourcelimits{resource="memory"}
# CPU throttling ratio
rate(containercpucfsthrottledperiodstotal[5m])
/ rate(containercpucfsperiodstotal[5m])
# CPU usage against request
rate(containercpuusagesecondstotal[5m])
/ on(namespace,pod,container)
kubepodcontainerresourcerequests{resource="cpu"}
5. Pending Pods: Scheduling Failures
A Pod that stays Pending has not been placed on any node. The event from the scheduler lists, for each reason, how many nodes were rejected.
kubectl describe pod <pod> -n <ns> | sed -n '/Events:/,$p'
# Warning FailedScheduling 0/5 nodes are available:
# 3 Insufficient cpu, 2 Insufficient memory.
# Requests committed vs allocatable, per node
kubectl describe node <node> | sed -n '/Allocated resources:/,/Events:/p'
# Actual usage per node
kubectl top node
Remember that the scheduler subtracts requests, not usage, from allocatable capacity, so the first comparison is always requested versus allocatable, and only then requested versus used.
6. CPU Issues
6.1 CPU throttling
Throttling is the most frequent hidden cause of latency. When a container reaches its CPU limit within a scheduling period (100 ms by default), the kernel pauses its threads until the next period. Average CPU can look far below the limit while short bursts are being throttled.
Symptoms: latency spikes, timeouts, slow start-up, failing probes, with moderate average CPU.
Evidence: a rising nrthrottled and throttledusec in cpu.stat, or a high throttled-periods ratio in monitoring.
Typical causes: limit too low for bursty or multi-threaded workloads; runtimes that create more threads than the limit allows.
Fixes: raise the CPU limit, or remove it deliberately if your platform policy allows (see the requests versus limits article), reduce thread and worker counts to match the limit, and make sure requests reflect real needs.
6.2 Node CPU contention
If many Pods on a node want CPU at the same time, each gets a share proportional to its request. Pods with low requests receive little CPU under contention, even if they have no limit.
Symptoms: slowness across several Pods on the same node; high node CPU in kubectl top node; no throttling in the affected containers.
Causes: requests set too low, so nodes are overcommitted; a noisy neighbour; too many Pods per node.
Fixes: raise requests to realistic values, spread load, use Guaranteed QoS or priority for critical Pods, add nodes.
6.3 High CPU usage in the application
Sometimes the container really does need the CPU. Compare usage with the request and limit, then investigate the application: profiling, a hot loop, garbage collection pressure, retries storms, or a traffic increase. Scaling out with the HPA is appropriate if the work can be split.
6.4 Runtime-specific CPU pitfalls
Thread and worker counts based on node cores: some runtimes size their pools from the number of CPUs they see. Older Go versions do not respect CPU limits when setting GOMAXPROCS (newer versions improved this, so check your version), and Node.js and Python worker pools often use fixed defaults. Set pool sizes explicitly to match the CPU limit.
JVM: modern JVMs are container-aware for CPU and memory, but check garbage collector threads and JIT compiler threads against small limits.
Start-up bursts: many applications use much more CPU while starting; a low limit slows start-up and can make liveness or startup probes fail.
7. Memory Issues
Memory problems come in three shapes: steady growth (a leak), sudden spikes (a burst of work), and a limit that is simply too low for normal operation. The graph of the working set tells you which one you have.
7.1 What counts toward the limit
The kernel charges more than the application heap: native allocations, thread stacks, page cache that is in active use, kernel memory such as socket buffers, and tmpfs files. A runtime heap configured equal to the container limit will always be killed, because other memory comes on top. Leave headroom (for example configure a JVM with MaxRAMPercentage rather than a fixed heap equal to the limit) and compare the working set metric, not raw usage including reclaimable cache, with the limit.
7.2 Requests that are too low
A memory request far below real usage makes the Pod look cheap to the scheduler. Many such Pods end up on one node, which then runs out of memory, and the Pods that exceed their requests are the first to be evicted or killed. This looks like a random node problem but is really a request problem.
8. OOM Kills and Evictions
Three different events are often called "OOM". Use this table to decide which one you have.
8.1 Diagnosing a container OOM kill
kubectl describe pod <pod> -n <ns>
# Containers:
# app:
# Last State: Terminated
# Reason: OOMKilled
# Exit Code: 137
# Restart Count: 5
# Limits:
# memory: 256Mi
kubectl exec <pod> -n <ns> -- cat /sys/fs/cgroup/memory.events
# oomkill counter increases with each kill
The example output is illustrative. If the Pod has several containers, check each one: a sidecar may be the one that is killed. Also remember that exit code 137 means the process received SIGKILL; it is OOMKilled only if the Reason says so. Other causes of SIGKILL include a manual kill or a termination that exceeded its grace period.
8.2 Diagnosing evictions
kubectl get pods -n <ns> | grep Evicted
kubectl describe pod <evicted-pod> -n <ns>
# Status: Failed
# Reason: Evicted
# Message: The node was low on resource: memory. ...
kubectl describe node <node> | grep -A8 Conditions
# MemoryPressure True
Evicted Pods remain as Failed objects until cleaned up. Controllers create replacements, so the application usually recovers, but the cause remains. Look at the memory requests and QoS class of the evicted Pods versus the ones that survived, and at the node's total requested memory against allocatable. Disk-related evictions mention ephemeral storage or DiskPressure instead; the cause then is container logs, emptyDir data or image storage, and the fix involves ephemeral-storage requests and limits.
# Clean up evicted or failed Pods after investigating (review the list first)
kubectl get pods -A --field-selector=status.phase=Failed
kubectl delete pods -A --field-selector=status.phase=Failed
8.3 When only a child process dies
If the kernel kills a worker process but not the main one, the container keeps running and no OOMKilled reason appears; users just see errors or lost workers. Evidence is in the kernel log and in memory.events. Newer systems may be configured to kill the whole container together, which makes this easier to see; the behaviour depends on the Kubernetes version and runtime.
9. Quota and LimitRange Errors
Resource problems also appear as API errors when a Pod is created. These are not scheduling failures; the Pod never exists.
kubectl describe quota -n <ns>
kubectl describe limitrange -n <ns>
10. Probes and Resource Pressure
Resource shortages often show up as probe failures, which then cause restarts that look unrelated. A container starved of CPU responds slowly, a liveness probe times out, and the kubelet restarts the container, which adds start-up CPU load and makes things worse.
Check the Events for Liveness probe failed and Readiness probe failed messages, and see whether they coincide with CPU throttling or node CPU pressure.
Use a startup probe for slow starters so liveness does not run during start-up.
Give probes realistic timeouts and thresholds, and keep liveness checks light and independent of external dependencies.
Fix the underlying CPU shortage first; loosening probes only hides it.
11. Case Studies
These scenarios are composites, shown to illustrate the method rather than specific incidents.
12. Hands-On Lab
This lab reproduces CPU throttling and a memory OOM kill and shows where the evidence appears. Use a test cluster. The stress image is a commonly used community tool; any equivalent works. The cgroup paths assume cgroup v2.
Create a Pod that burns two CPU threads under a 200m limit.
Observe usage with kubectl top and read the throttling counters in cpu.stat.
Create a Pod whose memory use exceeds its limit.
Read the Last State, the exit code and memory.events.
cat <<EOF | kubectl apply -f -
apiVersion: v1
kind: Pod
metadata:
name: cpu-lab
spec:
containers:
- name: stress
image: polinux/stress
command: ["stress"]
args: ["--cpu", "2"]
resources:
requests:
cpu: 100m
limits:
cpu: 200m
---
apiVersion: v1
kind: Pod
metadata:
name: mem-lab
spec:
containers:
- name: stress
image: polinux/stress
command: ["stress"]
args: ["--vm", "1", "--vm-bytes", "300M", "--vm-hang", "1"]
resources:
requests:
memory: 100Mi
limits:
memory: 200Mi
EOF
# CPU: usage is held near the limit, throttling counters grow
kubectl top pod cpu-lab
kubectl exec cpu-lab -- cat /sys/fs/cgroup/cpu.stat
# Memory: restarts, OOMKilled, exit code 137
kubectl get pod mem-lab -w
kubectl describe pod mem-lab | grep -A6 "Last State"
# Clean up
kubectl delete pod cpu-lab mem-lab
Expected result (illustrative): for cpu-lab, usage stays near 200m even though two busy threads would use about two CPUs, and nrthrottled and throttled_usec increase. For mem-lab, the container is OOMKilled with exit code 137 and ends in CrashLoopBackOff. If the cgroup files are not found, your nodes may use cgroup v1; use the v1 paths described earlier. If the image has no cat, use a debug container.
13. Fix Catalogue
14. Prevention
Require requests and a memory limit on every production container, using LimitRange defaults and admission policies.
Right-size continuously with VPA recommendations or usage reviews.
Alert early: working set above 80 to 90 percent of the limit, throttling ratio above a threshold, Pending Pods for several minutes, MemoryPressure, OOMKilled reasons, restart rates.
Load test with realistic traffic before release, and compare behaviour against limits.
Choose QoS deliberately: Guaranteed for critical workloads, Burstable for most, no BestEffort in production.
Reserve system resources on nodes with kube-reserved and system-reserved, and watch node pressure.
Keep a runbook with the triage steps and commands from this article, so that on-call engineers do not start from zero.
15. Best Practices
Start every investigation with kubectl describe pod; it answers most questions.
Compare requests, limits and real usage side by side before changing anything.
Separate container-level problems (limits) from node-level problems (requests and capacity).
Use previous logs, events, cgroup counters and kernel logs as evidence; avoid guessing.
Change one thing at a time and verify with metrics.
Treat repeated OOM kills and throttling as design signals, not only sizing tasks.
Document findings and update dashboards and alerts so the next incident is shorter.
16. Conclusion
Resource troubleshooting in Kubernetes is manageable once you know where to look. Pending Pods are about requests and capacity. Slowness with calm CPU graphs is usually throttling. Restarts with exit code 137 are container OOM kills against a memory limit. Evictions are node-level pressure combined with low requests or low priority. Each family has its own evidence in events, container status, cgroup files, kernel logs and metrics. Follow the triage method, collect the evidence, apply the matching fix, and then invest in prevention through accurate requests, sensible limits, good QoS choices and early alerts. That turns resource incidents from mysteries into routine, well-understood work.