Kubernetes Resource Quotas and LimitRanges
Kubernetes Resource Quotas and LimitRanges Explained
bitcodematrix.com | Kubernetes Series
1. Introduction
A Kubernetes cluster is a shared pool of CPU, memory, storage and API objects. Without any limits, one team's runaway Deployment, a forgotten batch job or a single container with a memory leak can consume capacity that other teams depend on. Kubernetes gives you two namespace-level tools to prevent this: the ResourceQuota, which caps what a whole namespace may use, and the LimitRange, which sets defaults and boundaries for each individual container, Pod or volume claim.
The two are often mentioned together, and for good reason: they complement each other, and in practice, one frequently does not work without the other. This article explains how requests and limits work, what each object controls, how they are enforced at admission, how to design quotas for teams, how they interact with rollouts and autoscaling, and how to monitor and troubleshoot them. It builds on the earlier articles in this series on admission controllers, ephemeral storage and storage classes.
2. A Quick Refresher on Requests and Limits
Every container can declare two numbers for CPU and memory, and both quotas and limit ranges are built on them:
CPU is measured in cores, so 1 is one core and 100m (millicores) is a tenth of a core. Memory is measured in bytes, with suffixes such as Mi and Gi, which are binary multiples. As in the ephemeral storage article, take care that Mi and M differ by about five percent.
The combination of requests and limits also determines the Quality of Service (QoS) class of a Pod, which governs who is evicted first when a node runs short of resources:
Guaranteed: every container has CPU and memory requests equal to its limits.
Burstable: at least one container has a request or limit, but the Pod is not Guaranteed.
BestEffort: no container sets any CPU or memory request or limit. These Pods are the first to be evicted under pressure.
3. Two Tools for Two Different Questions
Figure 1: LimitRange and ResourceQuota are both checked when a Pod is created.
Both are namespaced objects and are enforced by admission controllers (LimitRanger and ResourceQuota), which means they apply when a request is made, not retroactively. Both are enabled by default in standard clusters.
4. LimitRange in Depth
4.1 An Example
apiVersion: v1
kind: LimitRange
metadata:
name: container-limits
namespace: shop
spec:
limits:
- type: Container
defaultRequest:
cpu: 100m
memory: 128Mi
default:
cpu: 500m
memory: 256Mi
min:
cpu: 50m
memory: 64Mi
max:
cpu: "2"
memory: 1Gi
maxLimitRequestRatio:
memory: "4"
- type: PersistentVolumeClaim
min:
storage: 1Gi
max:
storage: 100Gi
4.2 Fields and Types
Container rules are evaluated for each container, including init containers.
Pod rules (with min, max and maxLimitRequestRatio, but no defaults) are evaluated against the sum of all containers in the Pod.
PersistentVolumeClaim rules restrict the requested storage size of each claim.
4.3 How Defaults Are Applied
When a container does not specify values, the LimitRanger plugin fills them in during the mutating phase of admission. As a rule:
A container that sets neither a request nor a limit receives both the defaultRequest and the default from the LimitRange.
A container that sets a limit but no request is given a request equal to its limit. A container that sets a request but no limit receives the default limit.
If the LimitRange defines a default limit but no defaultRequest, the request follows the limit.
Defaults are written into the Pod specification when it is created. That has two consequences: you can see them with kubectl get pod -o yaml, and changing the LimitRange later does not modify existing Pods.
4.4 How Validation Works
After defaults are set, the validating phase checks each value against min, max and the ratio. A violation is rejected with a clear message:
Error from server (Forbidden): pods "big" is forbidden:
maximum memory usage per Container is 1Gi, but limit is 2Gi
4.5 Effects on Quality of Service
Because defaults add requests and limits, a Pod that would have been BestEffort becomes Burstable (or Guaranteed, when the defaults are equal). That is usually what you want, since it removes the BestEffort class from the namespace and makes evictions fairer and more predictable.
4.6 Gotchas
Existing Pods are unchanged when you create or modify a LimitRange. Only new Pods are affected.
Several LimitRanges in a namespace can conflict, and which default wins is not defined. Keep one LimitRange per type, or a single object with all rules.
Minimums can break system add-ons that run in the namespace with small requests, so apply LimitRanges where you control all workloads.
A very low default limit causes mysterious out-of-memory kills for applications that are not given explicit values. Choose defaults based on measurement, and communicate them.
5. ResourceQuota in Depth
5.1 An Example
apiVersion: v1
kind: ResourceQuota
metadata:
name: team-quota
namespace: shop
spec:
hard:
requests.cpu: "10"
requests.memory: 20Gi
limits.memory: 40Gi
requests.ephemeral-storage: 50Gi
requests.storage: 500Gi
persistentvolumeclaims: "20"
pods: "50"
services: "20"
services.loadbalancers: "2"
count/deployments.apps: "30"
count/secrets: "100"
5.2 What Can Be Limited
5.3 How Enforcement Works
At admission. When a request would make the namespace's usage exceed any hard limit, the ResourceQuota plugin rejects it with 403 Forbidden. The message names the quota, what was requested, the current use and the limit.
Totals use the Pod's effective resources. Requests are summed across the containers of the Pod, with the rules for init containers applied. Pods in a terminal state (Succeeded or Failed) no longer count toward compute quota.
Compute quotas require explicit values. If a quota covers CPU or memory, every new Pod must specify the matching request or limit, or it is rejected with a message such as "must specify limits.memory". A LimitRange with defaults solves this, which is the main reason to use both.
Quotas do not evict. Lowering a quota below current use does not remove anything. It only blocks new requests until usage falls.
Quotas are not reservations. They set a ceiling. The sum of all quotas can exceed the real capacity of the cluster, which is common and called overcommitting. A namespace can be within its quota and still fail to get its Pods scheduled, if the cluster is full.
Usage is tracked by a controller. The status of the quota shows the hard limits and the current use. Changes to usage are reflected quickly, but may lag by a short time, particularly after you edit the quota itself.
A rejection looks like this:
Error from server (Forbidden): pods "api-7d9f" is forbidden:
exceeded quota: team-quota, requested: requests.cpu=500m,
used: requests.cpu=9800m, limited: requests.cpu=10
5.4 Quota Scopes
A quota can apply only to a subset of Pods, by adding scopes. Common scopes are:
For example, this quota limits the resources used by Pods of a high priority class, so that a team cannot make everything important:
apiVersion: v1
kind: ResourceQuota
metadata:
name: high-priority-quota
namespace: shop
spec:
hard:
cpu: "20"
memory: 40Gi
pods: "10"
scopeSelector:
matchExpressions:
- operator: In
scopeName: PriorityClass
values: ["high"]
Note that some priority-class quota setups also require the administrator to control which namespaces may use a given class, so check how your cluster handles that.
6. How They Work Together
The LimitRanger and the ResourceQuota run in the same admission chain, one after the other, as the diagram in Section 3 showed. The defaults come first, so that the quota sees complete values:
You create a Pod, and some containers omit their requests and limits.
The LimitRanger adds the default requests and limits, and rejects the Pod if any value is below the minimum or above the maximum.
The ResourceQuota plugin adds the Pod's effective requests and limits to the namespace's current usage, and rejects the Pod if any hard limit would be exceeded.
If both pass, the Pod is stored, and the quota's used values are updated.
The practical conclusion: deploy both in every shared namespace. The LimitRange makes quota enforcement painless for developers, and the quota makes the LimitRange meaningful for the cluster.
7. Designing Quotas for Teams
7.1 Choose What to Limit
Always: requests.cpu, requests.memory, the number of Pods, persistent storage and the number of claims.
Often: limits.memory, ephemeral storage, LoadBalancer Services (which cost money and public addresses), NodePort Services and the count of Secrets and ConfigMaps (to protect the API server and etcd).
Carefully: limits.cpu. A quota on CPU limits forces every container to define a CPU limit, which many teams avoid because CPU limits cause throttling even when the node has idle capacity. A quota on requests.cpu governs the capacity that is reserved, without that side effect.
7.2 Size by Tiers
Rather than negotiating each number, offer a few standard sizes, and let teams request a tier. The values below are only illustrations:
Base the numbers on real data: the observed usage of existing workloads, the expected growth, and the rollout and autoscaling behaviour described in Section 8. Review them regularly. Quotas that are too tight create friction and workarounds, and those that are too generous achieve nothing.
7.3 A Namespace Template
Combine the controls into one reusable bundle that is applied whenever a namespace is created, for example by a GitOps repository or an operator. It also includes the Pod Security labels from the previous articles:
apiVersion: v1
kind: Namespace
metadata:
name: team-shop
labels:
pod-security.kubernetes.io/enforce: restricted
pod-security.kubernetes.io/warn: restricted
---
apiVersion: v1
kind: ResourceQuota
metadata:
name: compute-and-objects
namespace: team-shop
spec:
hard:
requests.cpu: "8"
requests.memory: 16Gi
limits.memory: 32Gi
requests.storage: 500Gi
persistentvolumeclaims: "20"
pods: "60"
services.loadbalancers: "1"
---
apiVersion: v1
kind: LimitRange
metadata:
name: defaults
namespace: team-shop
spec:
limits:
- type: Container
defaultRequest:
cpu: 100m
memory: 128Mi
default:
memory: 256Mi
max:
cpu: "4"
memory: 8Gi
Notice that the LimitRange gives a default memory limit but no default CPU limit. The quota counts requests.cpu, which the defaultRequest provides, and limits.memory, which the default provides, so every Pod passes through without the team needing to know the details.
7.4 Different Environments
Development and test: smaller quotas, BestEffort allowed or capped at a low number, and frequent clean-up of unused resources.
Production: quotas based on capacity planning, headroom for rollouts and scaling, and no BestEffort Pods, enforced with a BestEffort-scoped quota of zero Pods.
Shared platform namespaces: quotas per tool, so that a monitoring or CI stack cannot starve applications.
8. Quotas, Rollouts and Autoscaling
Quotas count what exists at this moment, including the temporary extra Pods that rollouts and scaling create. Forgetting this is the most common reason for confusing failures.
9. Monitoring and Alerting
Quotas fail loudly at the worst moment, when a rollout or an autoscaling event needs capacity, so watch the utilisation instead of waiting for errors. Inspect them directly:
kubectl get resourcequota -A
kubectl describe resourcequota team-quota -n shop
kubectl describe limitrange -n shop
The describe output for a quota prints a table of each resource with its used and hard values. For continuous monitoring, kube-state-metrics exposes the data through the kuberesourcequota metric, with labels for the quota, the namespace, the resource and the type, which is either hard or used. An alert for namespaces that are close to a limit can be written as:
kuberesourcequota{type="used"}
/ ignoring(type) kube_resourcequota{type="hard"} > 0.9
Combine it with alerts on ReplicaSet events that mention exceeded quota, and with a dashboard that compares each namespace's used quota, requests and real consumption. The comparison between requests and actual usage reveals over-provisioned workloads, which is where most quota pressure comes from.
10. Hands-On Walkthrough
Create a namespace with a LimitRange that defines defaults and a maximum.
kubectl create namespace quota-demo
kubectl apply -n quota-demo -f - <<'EOF'
apiVersion: v1
kind: LimitRange
metadata:
name: defaults
spec:
limits:
- type: Container
defaultRequest:
cpu: 100m
memory: 64Mi
default:
cpu: 200m
memory: 128Mi
max:
cpu: "1"
memory: 512Mi
EOF
Add a ResourceQuota that covers requests and limits for memory and CPU, and the number of Pods.
kubectl apply -n quota-demo -f - <<'EOF'
apiVersion: v1
kind: ResourceQuota
metadata:
name: compute-quota
spec:
hard:
requests.cpu: "1"
requests.memory: 512Mi
limits.memory: 1Gi
pods: "5"
EOF
Create a Pod without any resources, and see that the defaults were injected, and that the quota counts them.
kubectl run plain -n quota-demo --image=busybox:1.36 \
--restart=Never --command -- sleep 3600
kubectl get pod plain -n quota-demo \
-o jsonpath='{.spec.containers[0].resources}'
kubectl describe resourcequota compute-quota -n quota-demo
The Pod shows the default request and limit from the LimitRange, and the quota's used column now includes 100m CPU and 64Mi memory.
Try a Pod that breaks the LimitRange maximum.
kubectl run big -n quota-demo --image=busybox:1.36 --restart=Never \
--overrides='{"spec":{"containers":[{"name":"big","image":"busybox:1.36","command":["sleep","3600"],"resources":{"limits":{"memory":"2Gi"}}}]}}'
The request is rejected with a message that the maximum memory usage per container is 512Mi, even though the quota still has room.
Exhaust the quota with a Deployment whose Pods request 250m of CPU each.
kubectl create deployment web -n quota-demo --image=busybox:1.36 \
--replicas=6 -- sleep 3600
kubectl set resources deployment web -n quota-demo \
--requests=cpu=250m,memory=64Mi --limits=memory=128Mi
kubectl get deployment web -n quota-demo
kubectl describe replicaset -n quota-demo | grep -i -B1 -A2 "exceeded quota"
Only some of the six replicas become ready, because the namespace has 1 CPU of requests in total, which is already partly used by the first Pod. The ReplicaSet's events contain "exceeded quota" messages, even though the Deployment itself was accepted without any complaint.
Free capacity by scaling down, and watch the missing Pods appear.
kubectl delete pod plain -n quota-demo
kubectl get pods -n quota-demo
kubectl describe resourcequota compute-quota -n quota-demo
Clean up.
kubectl delete namespace quota-demo
11. Troubleshooting
12. Best Practices
Deploy a ResourceQuota and a LimitRange together in every shared namespace, and create them automatically with the namespace.
Use the LimitRange to supply sensible defaults and sanity limits, so that developers do not need to know about the quota to deploy.
Prefer quotas on requests for CPU, and be careful with quotas on CPU limits, because they force limits that can throttle applications.
Quota memory limits as well as requests, so that one namespace cannot over-commit node memory dangerously.
Cap object counts that stress the control plane or cost money, such as Services of type LoadBalancer, Secrets, ConfigMaps and claims.
Offer a few standard tiers, base them on measured usage, and review them regularly.
Leave headroom for rolling updates, autoscaling and Jobs, and keep HPA maximums consistent with the quota.
Prevent BestEffort Pods in production with a BestEffort-scoped quota, or with defaults that make every Pod Burstable at least.
Remember that quotas are ceilings, not reservations, and track real cluster capacity separately.
Alert on quota utilisation above about 80 or 90 percent, and on quota errors in ReplicaSet events.
Keep quota and LimitRange manifests in version control, with the change process for increases clearly documented.
Combine them with Pod Security labels, RBAC and network policies in the same namespace template.
13. Conclusion
ResourceQuotas and LimitRanges turn a shared Kubernetes cluster from a free-for-all into a managed environment. The LimitRange looks at each container, Pod and claim: it fills in defaults for the forgetful and rejects values that are unreasonably small or large. The ResourceQuota looks at the whole namespace: it caps total resources and counts of objects, so that no team can take more than its share. Both act at admission time, which makes them predictable, and both are best used together.
Design them as part of the namespace template, size them from real data, leave room for rollouts and autoscaling, monitor their utilisation and treat increases as a deliberate process. With these habits in place, capacity problems show up as early, readable signals, instead of as outages that affect other teams.