kubernetes

Kubernetes Resource Troubleshooting: CPU, Memory and OOM Issues

By Shubhankar Tripathi • • 5 min read

Kubernetes Resource Troubleshooting: CPU, Memory and OOM Issues

bitcodematrix.com | Kubernetes Series

1. Introduction

Resource problems are among the most common and most confusing issues in Kubernetes. A Pod stays Pending although the cluster looks idle. A container restarts every few minutes with no error in its logs. A service becomes slow while its CPU graph looks calm. A node starts evicting Pods at night. The symptoms differ, but the causes all come back to the same few things: requests, limits, node capacity and how the Linux kernel enforces them.

This article is a practical troubleshooting playbook. It gives you a triage method, a toolbox of commands, and a walk through the four main symptom families: Pending Pods, CPU problems, memory problems and OOM kills, and evictions. It finishes with short case studies, a hands-on lab, a fix catalogue and a prevention checklist. It is the troubleshooting companion to the earlier articles on requests and limits, QoS classes, OOMKilled, scheduling and autoscaling.

After reading it you should be able to:

  • Classify a resource problem from its symptom within a minute or two.

  • Use kubectl, metrics and cgroup files to find the root cause.

  • Tell CPU throttling, node CPU starvation and application slowness apart.

  • Distinguish a container OOM kill, a node-level OOM and an eviction.

  • Apply the right short-term and long-term fix, and prevent recurrence.

2. A Quick Mental Model

Most resource troubleshooting becomes easier with five facts in mind.

Fact

Consequence

The scheduler uses requests, not usage

Pods can be Pending on a cluster whose real usage is low, and nodes can be overloaded while requests look fine

CPU is compressible

Exceeding a CPU limit causes throttling (slowness), not a kill

Memory is incompressible

Exceeding a memory limit causes an OOM kill; node memory shortage causes eviction or a node-level OOM

Limits are enforced by the kernel through cgroups

Evidence is available in cgroup files and kernel logs, not only in Kubernetes objects

QoS class and usage relative to requests decide who suffers under node pressure

BestEffort and Burstable Pods above their requests are affected first


3. A Triage Method

Four symptom families: Pending, Restarting, Slow and Evicted, each with the first things to check

Figure 1: Triage by symptom

Ask these questions in order:

  1. What is the Pod's state? Pending, Running with restarts, Failed or Evicted, or Running but slow.

  2. What do the events and last state say? kubectl describe pod answers most questions: scheduling messages, Last State, Reason, Exit Code.

  3. Is it one container or the whole node? Check whether other Pods on the same node are affected, and whether the node reports pressure conditions.

  4. What do the numbers say? Compare requests, limits and actual usage (kubectl top and metrics).

  5. What changed? New release, new traffic pattern, new limits, node upgrade, new neighbour workloads.

4. The Troubleshooting Toolbox

4.1 Kubernetes level

# State, restarts, node

kubectl get pod <pod> -n <ns> -o wide

 

# Events, Last State, Reason, Exit Code, requests and limits

kubectl describe pod <pod> -n <ns>

 

# Previous container instance logs (after a crash)

kubectl logs <pod> -n <ns> --previous

 

# Current usage (needs metrics-server)

kubectl top pod <pod> -n <ns> --containers

kubectl top node

 

# What the node has committed and its conditions

kubectl describe node <node>

 

# Recent events, newest last

kubectl get events -n <ns> --sort-by=.lastTimestamp

 

# QoS class and last termination reason

kubectl get pod <pod> -n <ns> -o jsonpath='{.status.qosClass}{"\n"}{.status.containerStatuses[*].lastState.terminated.reason}{"\n"}'


4.2 Inside the container (cgroup files)

Because the kernel enforces limits, the cgroup files show exactly what the container sees. The paths below are for cgroup v2, which most modern distributions use. On cgroup v1, the equivalents live under /sys/fs/cgroup/cpu and /sys/fs/cgroup/memory with different file names, such as memory.limitinbytes and memory.usageinbytes. The container image must contain a shell and cat for kubectl exec to work.

# CPU limit as "quota period" in microseconds, e.g. 50000 100000 = 0.5 CPU

kubectl exec <pod> -n <ns> -- cat /sys/fs/cgroup/cpu.max

 

# CPU usage and throttling counters

kubectl exec <pod> -n <ns> -- cat /sys/fs/cgroup/cpu.stat

#   usageusec, nrperiods, nrthrottled, throttledusec

 

# Memory limit, current usage and OOM counters

kubectl exec <pod> -n <ns> -- cat /sys/fs/cgroup/memory.max

kubectl exec <pod> -n <ns> -- cat /sys/fs/cgroup/memory.current

kubectl exec <pod> -n <ns> -- cat /sys/fs/cgroup/memory.events

#   oom, oomkill


If the image has no shell, use an ephemeral debug container that shares the target's process namespace, for example kubectl debug -it <pod> --image=busybox --target=<container>. Note that the cgroup files seen from a debug container may differ depending on the runtime, so treat them as a guide.

4.3 Node level and monitoring

# Kernel log entries for OOM kills (on the node)

journalctl -k | grep -i -E "out of memory|killed process"

 

# Prometheus examples (verify metric names in your setup)

# Working set against the limit

containermemoryworkingsetbytes{container!=""}

  / on(namespace,pod,container)

  kubepodcontainerresourcelimits{resource="memory"}

 

# CPU throttling ratio

rate(containercpucfsthrottledperiodstotal[5m])

  / rate(containercpucfsperiodstotal[5m])

 

# CPU usage against request

rate(containercpuusagesecondstotal[5m])

  / on(namespace,pod,container)

  kubepodcontainerresourcerequests{resource="cpu"}


5. Pending Pods: Scheduling Failures

A Pod that stays Pending has not been placed on any node. The event from the scheduler lists, for each reason, how many nodes were rejected.

kubectl describe pod <pod> -n <ns> | sed -n '/Events:/,$p'

# Warning  FailedScheduling  0/5 nodes are available:

#   3 Insufficient cpu, 2 Insufficient memory.


Cause

How to recognise it

Fix

Requests larger than free capacity on every node

Insufficient cpu or memory; kubectl describe node shows Allocated resources near 100% of requests

Reduce requests if inflated; add nodes; enable node autoscaling

Request larger than any single node

Insufficient on all nodes even when the cluster is empty

Lower the request or add a larger node type

Fragmentation

Plenty of total free CPU, but split across many nodes

Smaller Pods, bin-packing scoring, add larger nodes

Requests far above real usage

Cluster looks full on paper, but kubectl top shows low usage

Right-size using usage data and VPA recommendations

Namespace quota exceeded

Pod rejected at creation, or Deployment shows failed create events

Free or raise the quota; see section 9

Other constraints

Events mention taints, affinity, node selector or volumes

Fix those constraints; adding capacity will not help

Too many Pods per node

Event says Too many pods

Add nodes or raise the per-node limit where supported


# Requests committed vs allocatable, per node

kubectl describe node <node> | sed -n '/Allocated resources:/,/Events:/p'

 

# Actual usage per node

kubectl top node


Remember that the scheduler subtracts requests, not usage, from allocatable capacity, so the first comparison is always requested versus allocatable, and only then requested versus used.

6. CPU Issues

6.1 CPU throttling

Throttling is the most frequent hidden cause of latency. When a container reaches its CPU limit within a scheduling period (100 ms by default), the kernel pauses its threads until the next period. Average CPU can look far below the limit while short bursts are being throttled.

  • Symptoms: latency spikes, timeouts, slow start-up, failing probes, with moderate average CPU.

  • Evidence: a rising nrthrottled and throttledusec in cpu.stat, or a high throttled-periods ratio in monitoring.

  • Typical causes: limit too low for bursty or multi-threaded workloads; runtimes that create more threads than the limit allows.

  • Fixes: raise the CPU limit, or remove it deliberately if your platform policy allows (see the requests versus limits article), reduce thread and worker counts to match the limit, and make sure requests reflect real needs.

6.2 Node CPU contention

If many Pods on a node want CPU at the same time, each gets a share proportional to its request. Pods with low requests receive little CPU under contention, even if they have no limit.

  • Symptoms: slowness across several Pods on the same node; high node CPU in kubectl top node; no throttling in the affected containers.

  • Causes: requests set too low, so nodes are overcommitted; a noisy neighbour; too many Pods per node.

  • Fixes: raise requests to realistic values, spread load, use Guaranteed QoS or priority for critical Pods, add nodes.

6.3 High CPU usage in the application

Sometimes the container really does need the CPU. Compare usage with the request and limit, then investigate the application: profiling, a hot loop, garbage collection pressure, retries storms, or a traffic increase. Scaling out with the HPA is appropriate if the work can be split.

6.4 Runtime-specific CPU pitfalls

  • Thread and worker counts based on node cores: some runtimes size their pools from the number of CPUs they see. Older Go versions do not respect CPU limits when setting GOMAXPROCS (newer versions improved this, so check your version), and Node.js and Python worker pools often use fixed defaults. Set pool sizes explicitly to match the CPU limit.

  • JVM: modern JVMs are container-aware for CPU and memory, but check garbage collector threads and JIT compiler threads against small limits.

  • Start-up bursts: many applications use much more CPU while starting; a low limit slows start-up and can make liveness or startup probes fail.

7. Memory Issues

Memory problems come in three shapes: steady growth (a leak), sudden spikes (a burst of work), and a limit that is simply too low for normal operation. The graph of the working set tells you which one you have.

Pattern in the graph

Likely cause

Next step

Steady climb until restart, repeating

Memory leak or unbounded cache

Heap or memory profile; bound caches; fix leak

Flat, then a sudden jump at certain times

Large request, import, report or batch

Stream data; limit concurrency; raise limit for the peak

Sits near the limit from the start

Limit too low for normal operation

Raise limit and request to measured peak plus headroom

Rises after load, never falls

Caching, fragmentation, or runtime not returning memory

Tune the runtime; accept a higher steady level; set the limit accordingly

Jumps when files are written to a volume

Memory-backed emptyDir or tmpfs counts as memory

Use disk-backed storage or a sizeLimit


7.1 What counts toward the limit

The kernel charges more than the application heap: native allocations, thread stacks, page cache that is in active use, kernel memory such as socket buffers, and tmpfs files. A runtime heap configured equal to the container limit will always be killed, because other memory comes on top. Leave headroom (for example configure a JVM with MaxRAMPercentage rather than a fixed heap equal to the limit) and compare the working set metric, not raw usage including reclaimable cache, with the limit.

7.2 Requests that are too low

A memory request far below real usage makes the Pod look cheap to the scheduler. Many such Pods end up on one node, which then runs out of memory, and the Pods that exceed their requests are the first to be evicted or killed. This looks like a random node problem but is really a request problem.

8. OOM Kills and Evictions

Three different events are often called "OOM". Use this table to decide which one you have.

Event

How to recognise it

Scope

Main fix

Container OOM kill

Last State Terminated, Reason OOMKilled, Exit Code 137; container restarts in place

One container

Raise limit, fix leak, tune runtime

Node-pressure eviction

Pod Failed with Reason Evicted; message "The node was low on resource: memory"; node condition MemoryPressure

Pods on one node, chosen by the kubelet

Raise requests, add capacity, improve QoS and priority

Node (system) OOM

Kernel log shows a process killed outside cgroup limits; node event such as SystemOOM; may hit system daemons

Anything on the node

Reserve system resources, add requests and limits, add capacity


8.1 Diagnosing a container OOM kill

kubectl describe pod <pod> -n <ns>

# Containers:

#   app:

#     Last State:  Terminated

#       Reason:    OOMKilled

#       Exit Code: 137

#     Restart Count: 5

#     Limits:

#       memory:    256Mi

 

kubectl exec <pod> -n <ns> -- cat /sys/fs/cgroup/memory.events

# oomkill counter increases with each kill


The example output is illustrative. If the Pod has several containers, check each one: a sidecar may be the one that is killed. Also remember that exit code 137 means the process received SIGKILL; it is OOMKilled only if the Reason says so. Other causes of SIGKILL include a manual kill or a termination that exceeded its grace period.

8.2 Diagnosing evictions

kubectl get pods -n <ns> | grep Evicted

kubectl describe pod <evicted-pod> -n <ns>

# Status:   Failed

# Reason:   Evicted

# Message:  The node was low on resource: memory. ...

 

kubectl describe node <node> | grep -A8 Conditions

# MemoryPressure   True


Evicted Pods remain as Failed objects until cleaned up. Controllers create replacements, so the application usually recovers, but the cause remains. Look at the memory requests and QoS class of the evicted Pods versus the ones that survived, and at the node's total requested memory against allocatable. Disk-related evictions mention ephemeral storage or DiskPressure instead; the cause then is container logs, emptyDir data or image storage, and the fix involves ephemeral-storage requests and limits.

# Clean up evicted or failed Pods after investigating (review the list first)

kubectl get pods -A --field-selector=status.phase=Failed

kubectl delete pods -A --field-selector=status.phase=Failed


8.3 When only a child process dies

If the kernel kills a worker process but not the main one, the container keeps running and no OOMKilled reason appears; users just see errors or lost workers. Evidence is in the kernel log and in memory.events. Newer systems may be configured to kill the whole container together, which makes this easier to see; the behaviour depends on the Kubernetes version and runtime.

9. Quota and LimitRange Errors

Resource problems also appear as API errors when a Pod is created. These are not scheduling failures; the Pod never exists.

Message (typical)

Cause

Fix

forbidden: exceeded quota: <name>, requested: requests.cpu=..., used: ..., limited: ...

ResourceQuota would be exceeded

kubectl describe quota; free or raise quota; reduce requests

must specify limits.memory / requests.cpu

Quota requires values and none were given

Add resources or a LimitRange with defaults

maximum cpu usage per Container is X, but limit is Y

LimitRange maximum exceeded

Lower the value or change the LimitRange

minimum memory usage per Container is X, but request is Y

LimitRange minimum not met

Raise the request


kubectl describe quota -n <ns>

kubectl describe limitrange -n <ns>


10. Probes and Resource Pressure

Resource shortages often show up as probe failures, which then cause restarts that look unrelated. A container starved of CPU responds slowly, a liveness probe times out, and the kubelet restarts the container, which adds start-up CPU load and makes things worse.

  • Check the Events for Liveness probe failed and Readiness probe failed messages, and see whether they coincide with CPU throttling or node CPU pressure.

  • Use a startup probe for slow starters so liveness does not run during start-up.

  • Give probes realistic timeouts and thresholds, and keep liveness checks light and independent of external dependencies.

  • Fix the underlying CPU shortage first; loosening probes only hides it.

11. Case Studies

Scenario

Findings

Resolution

Checkout API has latency spikes every few minutes

Average CPU 40% of limit, but throttled-periods ratio is high during spikes; CPU limit 500m with a multi-threaded runtime

Raised the CPU limit, capped worker threads, kept request at realistic usage; spikes disappeared

Java service OOMKilled under load with limit 1Gi

Heap fixed at 1g with -Xmx; metaspace, thread stacks and native memory add on top

Set heap to a percentage of the container limit and raised the limit to 1.5Gi; kills stopped

Pods Pending although the cluster is mostly idle

Each Pod requests 4 CPU but uses 0.3; node allocation near 100% of requests

Right-sized requests with VPA recommendations; the Cluster Autoscaler stopped adding unused nodes

Evictions every night on one node

Batch Jobs without requests (BestEffort) run at night and push the node into MemoryPressure

Added requests and limits to the jobs, moved them to their own node group, set a LimitRange default


These scenarios are composites, shown to illustrate the method rather than specific incidents.

12. Hands-On Lab

This lab reproduces CPU throttling and a memory OOM kill and shows where the evidence appears. Use a test cluster. The stress image is a commonly used community tool; any equivalent works. The cgroup paths assume cgroup v2.

  1. Create a Pod that burns two CPU threads under a 200m limit.

  2. Observe usage with kubectl top and read the throttling counters in cpu.stat.

  3. Create a Pod whose memory use exceeds its limit.

  4. Read the Last State, the exit code and memory.events.

cat <<EOF | kubectl apply -f -

apiVersion: v1

kind: Pod

metadata:

  name: cpu-lab

spec:

  containers:

    - name: stress

      image: polinux/stress

      command: ["stress"]

      args: ["--cpu", "2"]

      resources:

        requests:

          cpu: 100m

        limits:

          cpu: 200m

---

apiVersion: v1

kind: Pod

metadata:

  name: mem-lab

spec:

  containers:

    - name: stress

      image: polinux/stress

      command: ["stress"]

      args: ["--vm", "1", "--vm-bytes", "300M", "--vm-hang", "1"]

      resources:

        requests:

          memory: 100Mi

        limits:

          memory: 200Mi

EOF

 

# CPU: usage is held near the limit, throttling counters grow

kubectl top pod cpu-lab

kubectl exec cpu-lab -- cat /sys/fs/cgroup/cpu.stat

 

# Memory: restarts, OOMKilled, exit code 137

kubectl get pod mem-lab -w

kubectl describe pod mem-lab | grep -A6 "Last State"

 

# Clean up

kubectl delete pod cpu-lab mem-lab


Expected result (illustrative): for cpu-lab, usage stays near 200m even though two busy threads would use about two CPUs, and nrthrottled and throttled_usec increase. For mem-lab, the container is OOMKilled with exit code 137 and ends in CrashLoopBackOff. If the cgroup files are not found, your nodes may use cgroup v1; use the v1 paths described earlier. If the image has no cat, use a debug container.

13. Fix Catalogue

Problem

Quick fix

Long-term fix

Pod Pending: Insufficient resources

Reduce requests or add a node

Right-size requests; node autoscaling; capacity planning

Throttling and latency

Raise or remove the CPU limit

Tune thread counts; size requests and limits from load tests

Node CPU contention

Move or scale Pods; raise requests

Realistic requests everywhere; spread; Guaranteed QoS for critical Pods

Container OOMKilled

Raise the memory limit

Find the leak or burst; tune runtime memory; stream large data

Node-pressure evictions

Add capacity; raise requests of victims

Set requests on all Pods; ban BestEffort in production; reserve system resources

Quota errors

Raise or free the quota

Team budgets, LimitRange defaults, capacity reviews

Probe failures under load

Raise timeouts temporarily

Fix CPU shortage; startup probes; lighter checks


14. Prevention

  • Require requests and a memory limit on every production container, using LimitRange defaults and admission policies.

  • Right-size continuously with VPA recommendations or usage reviews.

  • Alert early: working set above 80 to 90 percent of the limit, throttling ratio above a threshold, Pending Pods for several minutes, MemoryPressure, OOMKilled reasons, restart rates.

  • Load test with realistic traffic before release, and compare behaviour against limits.

  • Choose QoS deliberately: Guaranteed for critical workloads, Burstable for most, no BestEffort in production.

  • Reserve system resources on nodes with kube-reserved and system-reserved, and watch node pressure.

  • Keep a runbook with the triage steps and commands from this article, so that on-call engineers do not start from zero.

15. Best Practices

  • Start every investigation with kubectl describe pod; it answers most questions.

  • Compare requests, limits and real usage side by side before changing anything.

  • Separate container-level problems (limits) from node-level problems (requests and capacity).

  • Use previous logs, events, cgroup counters and kernel logs as evidence; avoid guessing.

  • Change one thing at a time and verify with metrics.

  • Treat repeated OOM kills and throttling as design signals, not only sizing tasks.

  • Document findings and update dashboards and alerts so the next incident is shorter.

16. Conclusion

Resource troubleshooting in Kubernetes is manageable once you know where to look. Pending Pods are about requests and capacity. Slowness with calm CPU graphs is usually throttling. Restarts with exit code 137 are container OOM kills against a memory limit. Evictions are node-level pressure combined with low requests or low priority. Each family has its own evidence in events, container status, cgroup files, kernel logs and metrics. Follow the triage method, collect the evidence, apply the matching fix, and then invest in prevention through accurate requests, sensible limits, good QoS choices and early alerts. That turns resource incidents from mysteries into routine, well-understood work.