Kubernetes OOMKilled Explained: Why Kubernetes Terminates Containers
Kubernetes OOMKilled Explained: Why Kubernetes Terminates Containers
bitcodematrix.com | Kubernetes Series
1. Introduction
OOMKilled is one of the most common reasons a container restarts in Kubernetes, and one of the most misunderstood. The name comes from "out of memory". It means that a container used more memory than it was allowed to, and the Linux kernel ended it. The container does not get a chance to clean up or write a log line, because the signal that stops it cannot be caught.
This article explains exactly what OOMKilled means, how Kubernetes and Linux cooperate to produce it, how it differs from Pod eviction and from node-level out-of-memory events, what really counts as "memory" for a container, and how to diagnose and fix it systematically. It builds on the earlier articles on requests and limits and on QoS classes.
After reading it you should be able to:
Explain the chain of events from a memory spike to exit code 137.
Tell a container OOM kill from an eviction and a node OOM event.
Identify what memory the kernel counts against a container limit.
Diagnose an OOMKilled container with kubectl, logs, metrics and node logs.
Fix the root cause and prevent recurrence.
2. What OOMKilled Means
When a container is created, the kubelet and the container runtime place it in a Linux control group (cgroup). If you set a memory limit in the container spec, that value becomes the maximum amount of memory the cgroup may use. When the container tries to use more and the kernel cannot reclaim enough memory inside the cgroup, the kernel's out-of-memory (OOM) killer selects a process in that cgroup and sends it SIGKILL.
Kubernetes then observes that the container terminated and records the reason. You will see the following in the container status:
Reason: OOMKilled
Exit Code: 137, which is 128 plus 9, the number of the SIGKILL signal
Because SIGKILL cannot be trapped, the application has no chance to run shutdown hooks, flush buffers or log an error. This is why OOM kills often leave nothing at all in the application logs, which makes them confusing for the first time.
Figure 1: From memory spike to OOMKilled and restart
3. Three Different Things Called "Out of Memory"
Teams often use "OOM" for several distinct events. They look similar from the outside but have different causes and different fixes.
The difference matters. A container OOM kill is almost always about that container: its limit is too low or its memory use is too high. An eviction or a node OOM is about the node as a whole: Pods are collectively using more than the node can supply, which points to missing or unrealistic requests, overcommitment, or too little capacity. The article on QoS classes explains how Pods are ranked in the latter two cases.
4. What Counts as Container Memory
A common surprise is that the memory counted against the limit is more than the heap your application reports. The cgroup tracks all memory charged to the container, including:
Anonymous memory: heap, stacks, and memory allocated by libraries and native code.
Page cache: file data cached in memory. Most of this can be reclaimed under pressure, so it normally does not cause kills by itself, but dirty or actively used cache can contribute.
Kernel memory: socket buffers, kernel data structures and similar allocations attributed to the cgroup.
tmpfs and memory-backed volumes: files written to an emptyDir with medium set to Memory count against the container memory, as do files written to a tmpfs mount.
All processes in the container: the main process, child processes, sidecar threads and anything else in the cgroup.
For monitoring, the metric that best matches what the kubelet and the kernel care about is the working set (containermemoryworkingsetbytes), which is usage minus inactive file cache that can easily be reclaimed. Plain "usage" including cache can look alarming without being dangerous. Compare the working set against the limit, not against the node size.
4.1 Which process dies?
The kernel chooses a process inside the cgroup, typically the one using the most memory. If it kills the main process (PID 1 of the container), the container exits with exit code 137 and is reported as OOMKilled. If it kills a child process, the container may keep running with a failed worker. This can look like a mysterious application error instead of an OOMKilled event. On newer systems with cgroup v2, the kubelet may configure the cgroup so that all processes in the container are killed together, which gives more consistent behaviour. Verify the behaviour for your Kubernetes version and runtime.
5. What Happens After the Kill
Once the container exits, the kubelet follows the Pod's restartPolicy. For Deployments, StatefulSets and DaemonSets this is Always, so the container is restarted in place on the same node. Each restart increments the restart count. If the container keeps failing, the kubelet applies an exponential back-off delay (starting at about 10 seconds and increasing up to a maximum of about five minutes), and the Pod shows the status CrashLoopBackOff. Note that CrashLoopBackOff is a waiting state between restarts, and the actual cause is recorded in the last terminated state. For Jobs with restartPolicy OnFailure or Never, the behaviour follows the Job rules and may end with a failed Job.
This is why you often see a Pod in CrashLoopBackOff and must look one level deeper to find OOMKilled as the underlying reason.
6. Diagnosing an OOMKilled Container
6.1 Confirm the reason
kubectl get pod <pod> -n shop
# STATUS: CrashLoopBackOff or OOMKilled, RESTARTS increasing
kubectl describe pod <pod> -n shop
# Containers:
# app:
# Last State: Terminated
# Reason: OOMKilled
# Exit Code: 137
# Restart Count: 6
# Limits:
# memory: 256Mi
To read the reason directly for scripting:
kubectl get pod <pod> -n shop \
-o jsonpath='{.status.containerStatuses[*].lastState.terminated.reason}'
6.2 Look at the previous container's logs
kubectl logs <pod> -n shop --previous
The logs of the previous container instance may end abruptly, with no error. That sudden stop is itself a clue.
6.3 Check the metrics
Use the monitoring system to compare memory working set to the limit over time. Typical Prometheus queries, assuming cAdvisor and kube-state-metrics are installed, include:
# Working set as a fraction of the limit, per container
containermemoryworkingsetbytes{container!=""}
/ on(namespace,pod,container)
kubepodcontainerresourcelimits{resource="memory"}
# Containers whose last termination was an OOM kill
kubepodcontainerstatuslastterminatedreason{reason="OOMKilled"} == 1
# Rate of OOM events seen by cAdvisor
increase(containeroomeventstotal[1h])
The shape of the graph tells you the cause. A steady climb until the kill suggests a leak. A sudden spike suggests a burst of work such as a large request, batch or file load. A flat line that sits right at the limit suggests the limit is simply too low for normal operation.
6.4 Check the node
If you have access to the node, the kernel log records every OOM kill with details of the cgroup and the memory breakdown:
journalctl -k | grep -i "killed process"
# or
dmesg | grep -i "out of memory"
Also review node events and conditions with kubectl describe node. A MemoryPressure condition or a SystemOOM event indicates a node-level problem, not just one container.
7. Common Causes and Their Fixes
7.1 Runtime-specific tuning
Java: use -XX:MaxRAMPercentage (for example 70 to 75) so heap, metaspace, thread stacks and native memory fit in the limit.
Go: set GOMEMLIMIT slightly below the container limit so the garbage collector works harder before the kernel intervenes.
Node.js: set --max-old-space-size below the limit, leaving room for buffers and native memory.
Python: watch for large in-memory data frames and caches; consider processing data in chunks.
7.2 Choosing a new limit
Do not simply double the limit and move on. Measure the working set over a representative period, identify the peak, and set the limit above it with a safety margin (many teams use 20 to 30 percent as a starting point, adjusted to their risk tolerance). Set the memory request close to typical usage, or equal to the limit for critical workloads, as discussed in the QoS article. Consider running the Vertical Pod Autoscaler in recommendation mode to validate your numbers.
8. Hands-On Walkthrough
This exercise reproduces an OOM kill and walks through the diagnosis. Use a test cluster. The stress image is a commonly used community tool; any equivalent image works.
Create a Pod whose container tries to allocate more memory than its limit.
Watch the Pod restart and move into CrashLoopBackOff.
Confirm the reason and exit code with describe.
Fix the problem by raising the limit, then recreate the Pod.
cat <<EOF | kubectl apply -f -
apiVersion: v1
kind: Pod
metadata:
name: oom-demo
spec:
containers:
- name: stress
image: polinux/stress
command: ["stress"]
args: ["--vm", "1", "--vm-bytes", "300M", "--vm-hang", "1"]
resources:
requests:
memory: 100Mi
limits:
memory: 200Mi
EOF
kubectl get pod oom-demo -w
kubectl describe pod oom-demo | grep -A6 "Last State"
Expected result: the Pod goes through Running, then OOMKilled, then CrashLoopBackOff as restarts accumulate. Last State shows Terminated with Reason OOMKilled and Exit Code 137. Now fix it:
kubectl delete pod oom-demo
# Edit the manifest: set limits.memory to 400Mi, then
kubectl apply -f oom-demo.yaml
kubectl get pod oom-demo
# The Pod stays Running; the container no longer exceeds its limit
The exact output formatting differs between kubectl versions. Clean up with kubectl delete pod oom-demo when finished.
9. Preventing OOM Kills
Right-size from data. Base requests and limits on measured working-set usage, including peaks.
Alert early. Alert when the working set exceeds roughly 80 to 90 percent of the limit for a sustained period, not only after a kill.
Alert on restarts and reasons. Watch kubepodcontainerstatuslastterminated_reason and restart counts.
Load test. Test with realistic and peak workloads before release.
Bound the application. Limit request sizes, concurrency, cache sizes and queue depths inside the application itself.
Use LimitRanges and quotas. Provide sensible defaults so that no container runs without a memory limit.
Plan node capacity. Keep requests realistic so that nodes are not routinely overcommitted.
Use graceful designs. Because SIGKILL cannot be handled, make the application restart-safe: idempotent work, durable state and readiness probes.
10. Troubleshooting Reference
11. Best Practices
Always set a memory limit in production, so one container cannot take down a node.
Set memory requests that reflect real usage so scheduling and eviction decisions are accurate.
Leave headroom between the application's configured memory and the container limit.
Prefer the working-set metric when comparing usage to limits.
Treat repeated OOM kills as a bug signal, not just a sizing task; investigate leaks.
Check every container in the Pod, including sidecars and init containers.
Use disk-backed volumes for large temporary data instead of memory-backed ones.
Monitor, alert and review OOM events as part of regular operations.
12. Conclusion
OOMKilled is Kubernetes and Linux doing exactly what the memory limit asked for: when a container goes beyond its allowance and the kernel cannot reclaim enough, the kernel ends it with SIGKILL, and the kubelet restarts it according to policy. The key to dealing with it is to separate the three kinds of out-of-memory events, to measure the right memory metric, and to follow the evidence from kubectl describe through logs and metrics to the node. Fix the cause, whether that is a limit that is too low, a leak, a spike or a misconfigured runtime, and then add alerts so that the next time you hear about it before your users do.