kubernetes

Kubernetes Resource Quotas and LimitRanges

By Shubhankar Tripathi • • 5 min read

Kubernetes Resource Quotas and LimitRanges Explained

bitcodematrix.com | Kubernetes Series

1. Introduction

A Kubernetes cluster is a shared pool of CPU, memory, storage and API objects. Without any limits, one team's runaway Deployment, a forgotten batch job or a single container with a memory leak can consume capacity that other teams depend on. Kubernetes gives you two namespace-level tools to prevent this: the ResourceQuota, which caps what a whole namespace may use, and the LimitRange, which sets defaults and boundaries for each individual container, Pod or volume claim.

The two are often mentioned together, and for good reason: they complement each other, and in practice, one frequently does not work without the other. This article explains how requests and limits work, what each object controls, how they are enforced at admission, how to design quotas for teams, how they interact with rollouts and autoscaling, and how to monitor and troubleshoot them. It builds on the earlier articles in this series on admission controllers, ephemeral storage and storage classes.

2. A Quick Refresher on Requests and Limits

Every container can declare two numbers for CPU and memory, and both quotas and limit ranges are built on them:

Concept

Request

Limit

Meaning

The amount that the container is guaranteed, and that the scheduler reserves on a node

The maximum that the container may use

Used by

The scheduler, to place the Pod, and by the kubelet for fair sharing and eviction decisions

The kubelet and the container runtime, which enforce it

CPU when exceeded

Not applicable

The container is throttled; CPU is a compressible resource

Memory when exceeded

Not applicable

The container is killed with an out-of-memory error; memory is not compressible

 

CPU is measured in cores, so 1 is one core and 100m (millicores) is a tenth of a core. Memory is measured in bytes, with suffixes such as Mi and Gi, which are binary multiples. As in the ephemeral storage article, take care that Mi and M differ by about five percent.

The combination of requests and limits also determines the Quality of Service (QoS) class of a Pod, which governs who is evicted first when a node runs short of resources:

  • Guaranteed: every container has CPU and memory requests equal to its limits.

  • Burstable: at least one container has a request or limit, but the Pod is not Guaranteed.

  • BestEffort: no container sets any CPU or memory request or limit. These Pods are the first to be evicted under pressure.

3. Two Tools for Two Different Questions

A Pod create request passes the LimitRanger and then the ResourceQuota admission checks before it is stored

Figure 1: LimitRange and ResourceQuota are both checked when a Pod is created.

Aspect

LimitRange

ResourceQuota

Question answered

What defaults and size boundaries apply to each object?

How much may the namespace use in total?

Scope of a rule

A single container, Pod or PersistentVolumeClaim

The sum over everything in the namespace

Can set defaults

Yes, for container requests and limits

No

Can reject requests

Yes, for values below the minimum or above the maximum

Yes, for requests that would exceed the totals

Limits object counts

No

Yes, such as the number of Pods, Services or claims

Typical owner

Platform team, as a safe default for everyone

Platform team, per team or per environment

 

Both are namespaced objects and are enforced by admission controllers (LimitRanger and ResourceQuota), which means they apply when a request is made, not retroactively. Both are enabled by default in standard clusters.

4. LimitRange in Depth

4.1 An Example

apiVersion: v1

kind: LimitRange

metadata:

  name: container-limits

  namespace: shop

spec:

  limits:

  - type: Container

    defaultRequest:

      cpu: 100m

      memory: 128Mi

    default:

      cpu: 500m

      memory: 256Mi

    min:

      cpu: 50m

      memory: 64Mi

    max:

      cpu: "2"

      memory: 1Gi

    maxLimitRequestRatio:

      memory: "4"

  - type: PersistentVolumeClaim

    min:

      storage: 1Gi

    max:

      storage: 100Gi

4.2 Fields and Types

Field

Meaning

type

What the rule applies to: Container, Pod or PersistentVolumeClaim

defaultRequest

The request that a container receives if it does not set one (Container type only)

default

The limit that a container receives if it does not set one (Container type only)

min

The smallest allowed request or limit. A smaller value is rejected

max

The largest allowed limit (and request). A larger value is rejected

maxLimitRequestRatio

The largest allowed ratio of limit to request, which controls how much a container may burst above what it reserved

 

  • Container rules are evaluated for each container, including init containers.

  • Pod rules (with min, max and maxLimitRequestRatio, but no defaults) are evaluated against the sum of all containers in the Pod.

  • PersistentVolumeClaim rules restrict the requested storage size of each claim.

4.3 How Defaults Are Applied

When a container does not specify values, the LimitRanger plugin fills them in during the mutating phase of admission. As a rule:

  • A container that sets neither a request nor a limit receives both the defaultRequest and the default from the LimitRange.

  • A container that sets a limit but no request is given a request equal to its limit. A container that sets a request but no limit receives the default limit.

  • If the LimitRange defines a default limit but no defaultRequest, the request follows the limit.

Defaults are written into the Pod specification when it is created. That has two consequences: you can see them with kubectl get pod -o yaml, and changing the LimitRange later does not modify existing Pods.

4.4 How Validation Works

After defaults are set, the validating phase checks each value against min, max and the ratio. A violation is rejected with a clear message:

Error from server (Forbidden): pods "big" is forbidden:

maximum memory usage per Container is 1Gi, but limit is 2Gi

4.5 Effects on Quality of Service

Because defaults add requests and limits, a Pod that would have been BestEffort becomes Burstable (or Guaranteed, when the defaults are equal). That is usually what you want, since it removes the BestEffort class from the namespace and makes evictions fairer and more predictable.

4.6 Gotchas

  • Existing Pods are unchanged when you create or modify a LimitRange. Only new Pods are affected.

  • Several LimitRanges in a namespace can conflict, and which default wins is not defined. Keep one LimitRange per type, or a single object with all rules.

  • Minimums can break system add-ons that run in the namespace with small requests, so apply LimitRanges where you control all workloads.

  • A very low default limit causes mysterious out-of-memory kills for applications that are not given explicit values. Choose defaults based on measurement, and communicate them.

5. ResourceQuota in Depth

5.1 An Example

apiVersion: v1

kind: ResourceQuota

metadata:

  name: team-quota

  namespace: shop

spec:

  hard:

    requests.cpu: "10"

    requests.memory: 20Gi

    limits.memory: 40Gi

    requests.ephemeral-storage: 50Gi

    requests.storage: 500Gi

    persistentvolumeclaims: "20"

    pods: "50"

    services: "20"

    services.loadbalancers: "2"

    count/deployments.apps: "30"

    count/secrets: "100"

5.2 What Can Be Limited

Category

Examples

Compute resources

requests.cpu, requests.memory, limits.cpu, limits.memory. The plain names cpu and memory are aliases for the requests. Huge pages can also be limited

Ephemeral storage

requests.ephemeral-storage, limits.ephemeral-storage (see the ephemeral storage article)

Persistent storage

requests.storage for the total requested size, persistentvolumeclaims for the number of claims, and per-class forms such as fast-ssd.storageclass.storage.k8s.io/requests.storage and fast-ssd.storageclass.storage.k8s.io/persistentvolumeclaims

Object counts

The count/<resource>.<group> form works for namespaced resources, for example count/deployments.apps, count/jobs.batch and count/secrets. Older names such as pods, services, configmaps and secrets also exist

Service types

services.loadbalancers and services.nodeports limit expensive or exposed Service types

Extended resources

Resources such as GPUs, in the form requests.<resource name>, for example requests.nvidia.com/gpu. Only requests can be limited, not limits

 

5.3 How Enforcement Works

  • At admission. When a request would make the namespace's usage exceed any hard limit, the ResourceQuota plugin rejects it with 403 Forbidden. The message names the quota, what was requested, the current use and the limit.

  • Totals use the Pod's effective resources. Requests are summed across the containers of the Pod, with the rules for init containers applied. Pods in a terminal state (Succeeded or Failed) no longer count toward compute quota.

  • Compute quotas require explicit values. If a quota covers CPU or memory, every new Pod must specify the matching request or limit, or it is rejected with a message such as "must specify limits.memory". A LimitRange with defaults solves this, which is the main reason to use both.

  • Quotas do not evict. Lowering a quota below current use does not remove anything. It only blocks new requests until usage falls.

  • Quotas are not reservations. They set a ceiling. The sum of all quotas can exceed the real capacity of the cluster, which is common and called overcommitting. A namespace can be within its quota and still fail to get its Pods scheduled, if the cluster is full.

  • Usage is tracked by a controller. The status of the quota shows the hard limits and the current use. Changes to usage are reflected quickly, but may lag by a short time, particularly after you edit the quota itself.

A rejection looks like this:

Error from server (Forbidden): pods "api-7d9f" is forbidden:

exceeded quota: team-quota, requested: requests.cpu=500m,

used: requests.cpu=9800m, limited: requests.cpu=10

5.4 Quota Scopes

A quota can apply only to a subset of Pods, by adding scopes. Common scopes are:

Scope

Applies to

BestEffort and NotBestEffort

Pods with or without the BestEffort QoS class. The BestEffort scope can limit only the number of Pods, which lets you cap or forbid them

Terminating and NotTerminating

Pods that have an active deadline set (such as many Jobs), or not

PriorityClass

Pods of a given priority class, selected with a scopeSelector

CrossNamespacePodAffinity

Pods that use affinity terms that reach across namespaces, which allows an administrator to restrict that feature

 

For example, this quota limits the resources used by Pods of a high priority class, so that a team cannot make everything important:

apiVersion: v1

kind: ResourceQuota

metadata:

  name: high-priority-quota

  namespace: shop

spec:

  hard:

    cpu: "20"

    memory: 40Gi

    pods: "10"

  scopeSelector:

    matchExpressions:

    - operator: In

      scopeName: PriorityClass

      values: ["high"]

Note that some priority-class quota setups also require the administrator to control which namespaces may use a given class, so check how your cluster handles that.

6. How They Work Together

The LimitRanger and the ResourceQuota run in the same admission chain, one after the other, as the diagram in Section 3 showed. The defaults come first, so that the quota sees complete values:

  1. You create a Pod, and some containers omit their requests and limits.

  2. The LimitRanger adds the default requests and limits, and rejects the Pod if any value is below the minimum or above the maximum.

  3. The ResourceQuota plugin adds the Pod's effective requests and limits to the namespace's current usage, and rejects the Pod if any hard limit would be exceeded.

  4. If both pass, the Pod is stored, and the quota's used values are updated.

Situation

Result

Quota covers CPU and memory, no LimitRange, and a Pod sets no resources

The Pod is rejected, because the quota demands explicit values

Quota and LimitRange with defaults, and a Pod sets no resources

Defaults are applied, and the Pod is admitted if the totals fit

LimitRange only

Each object is bounded, but the namespace as a whole is not

Quota only, with all Pods setting resources

The namespace is bounded in total, but single Pods can be arbitrarily large, up to the quota

A Pod asks for more than max

LimitRange rejects it, even if the quota has space

A Pod is within max but the namespace is full

ResourceQuota rejects it

 

The practical conclusion: deploy both in every shared namespace. The LimitRange makes quota enforcement painless for developers, and the quota makes the LimitRange meaningful for the cluster.

7. Designing Quotas for Teams

7.1 Choose What to Limit

  • Always: requests.cpu, requests.memory, the number of Pods, persistent storage and the number of claims.

  • Often: limits.memory, ephemeral storage, LoadBalancer Services (which cost money and public addresses), NodePort Services and the count of Secrets and ConfigMaps (to protect the API server and etcd).

  • Carefully: limits.cpu. A quota on CPU limits forces every container to define a CPU limit, which many teams avoid because CPU limits cause throttling even when the node has idle capacity. A quota on requests.cpu governs the capacity that is reserved, without that side effect.

7.2 Size by Tiers

Rather than negotiating each number, offer a few standard sizes, and let teams request a tier. The values below are only illustrations:

Tier

requests.cpu

requests.memory

requests.storage

pods

claims

Small

2

4Gi

50Gi

20

5

Medium

8

16Gi

500Gi

60

20

Large

32

64Gi

2Ti

200

50

 

Base the numbers on real data: the observed usage of existing workloads, the expected growth, and the rollout and autoscaling behaviour described in Section 8. Review them regularly. Quotas that are too tight create friction and workarounds, and those that are too generous achieve nothing.

7.3 A Namespace Template

Combine the controls into one reusable bundle that is applied whenever a namespace is created, for example by a GitOps repository or an operator. It also includes the Pod Security labels from the previous articles:

apiVersion: v1

kind: Namespace

metadata:

  name: team-shop

  labels:

    pod-security.kubernetes.io/enforce: restricted

    pod-security.kubernetes.io/warn: restricted

---

apiVersion: v1

kind: ResourceQuota

metadata:

  name: compute-and-objects

  namespace: team-shop

spec:

  hard:

    requests.cpu: "8"

    requests.memory: 16Gi

    limits.memory: 32Gi

    requests.storage: 500Gi

    persistentvolumeclaims: "20"

    pods: "60"

    services.loadbalancers: "1"

---

apiVersion: v1

kind: LimitRange

metadata:

  name: defaults

  namespace: team-shop

spec:

  limits:

  - type: Container

    defaultRequest:

      cpu: 100m

      memory: 128Mi

    default:

      memory: 256Mi

    max:

      cpu: "4"

      memory: 8Gi

Notice that the LimitRange gives a default memory limit but no default CPU limit. The quota counts requests.cpu, which the defaultRequest provides, and limits.memory, which the default provides, so every Pod passes through without the team needing to know the details.

7.4 Different Environments

  • Development and test: smaller quotas, BestEffort allowed or capped at a low number, and frequent clean-up of unused resources.

  • Production: quotas based on capacity planning, headroom for rollouts and scaling, and no BestEffort Pods, enforced with a BestEffort-scoped quota of zero Pods.

  • Shared platform namespaces: quotas per tool, so that a monitoring or CI stack cannot starve applications.

8. Quotas, Rollouts and Autoscaling

Quotas count what exists at this moment, including the temporary extra Pods that rollouts and scaling create. Forgetting this is the most common reason for confusing failures.

Situation

What happens and what to do

Rolling update of a Deployment

The default strategy creates extra Pods (maxSurge) before it removes old ones. If the quota has no spare capacity for the surge, new Pods cannot be created and the rollout stalls. Keep headroom for at least the surge, or set maxSurge to 0 with maxUnavailable above 0

Horizontal Pod Autoscaler

HPA can raise the replica count only as far as the quota allows. When the quota is full, the ReplicaSet reports quota errors, and the HPA shows that it wants more replicas than it gets. Set maxReplicas consistently with the quota

Cluster autoscaler

It reacts to Pods that are pending because no node fits. A Pod blocked by quota is never created, so no node is added. Quota and cluster capacity are separate limits

Jobs and CronJobs

Parallel Jobs consume quota while running. Completed Pods stop counting for compute but remain as objects until cleaned up, and so can still count toward object-count quotas. Use ttlSecondsAfterFinished and history limits

Deployments and enforcement

A Deployment is accepted even if its Pods will exceed the quota. The failures appear as events on the ReplicaSet, in the same way as for Pod Security

PersistentVolumeClaims

The requested size counts toward the storage quota, even if the bound PV is larger. Released volumes under Retain do not count once the claim is deleted

StatefulSets

Each replica brings its own claim, so storage and claim-count quotas matter when scaling up

 

9. Monitoring and Alerting

Quotas fail loudly at the worst moment, when a rollout or an autoscaling event needs capacity, so watch the utilisation instead of waiting for errors. Inspect them directly:

kubectl get resourcequota -A

kubectl describe resourcequota team-quota -n shop

kubectl describe limitrange -n shop

The describe output for a quota prints a table of each resource with its used and hard values. For continuous monitoring, kube-state-metrics exposes the data through the kuberesourcequota metric, with labels for the quota, the namespace, the resource and the type, which is either hard or used. An alert for namespaces that are close to a limit can be written as:

kuberesourcequota{type="used"}

  / ignoring(type) kube_resourcequota{type="hard"} > 0.9

Combine it with alerts on ReplicaSet events that mention exceeded quota, and with a dashboard that compares each namespace's used quota, requests and real consumption. The comparison between requests and actual usage reveals over-provisioned workloads, which is where most quota pressure comes from.

10. Hands-On Walkthrough

  1. Create a namespace with a LimitRange that defines defaults and a maximum.

kubectl create namespace quota-demo

kubectl apply -n quota-demo -f - <<'EOF'

apiVersion: v1

kind: LimitRange

metadata:

  name: defaults

spec:

  limits:

  - type: Container

    defaultRequest:

      cpu: 100m

      memory: 64Mi

    default:

      cpu: 200m

      memory: 128Mi

    max:

      cpu: "1"

      memory: 512Mi

EOF

  1. Add a ResourceQuota that covers requests and limits for memory and CPU, and the number of Pods.

kubectl apply -n quota-demo -f - <<'EOF'

apiVersion: v1

kind: ResourceQuota

metadata:

  name: compute-quota

spec:

  hard:

    requests.cpu: "1"

    requests.memory: 512Mi

    limits.memory: 1Gi

    pods: "5"

EOF

  1. Create a Pod without any resources, and see that the defaults were injected, and that the quota counts them.

kubectl run plain -n quota-demo --image=busybox:1.36 \

  --restart=Never --command -- sleep 3600

kubectl get pod plain -n quota-demo \

  -o jsonpath='{.spec.containers[0].resources}'

kubectl describe resourcequota compute-quota -n quota-demo

The Pod shows the default request and limit from the LimitRange, and the quota's used column now includes 100m CPU and 64Mi memory.

  1. Try a Pod that breaks the LimitRange maximum.

kubectl run big -n quota-demo --image=busybox:1.36 --restart=Never \

  --overrides='{"spec":{"containers":[{"name":"big","image":"busybox:1.36","command":["sleep","3600"],"resources":{"limits":{"memory":"2Gi"}}}]}}'

The request is rejected with a message that the maximum memory usage per container is 512Mi, even though the quota still has room.

  1. Exhaust the quota with a Deployment whose Pods request 250m of CPU each.

kubectl create deployment web -n quota-demo --image=busybox:1.36 \

  --replicas=6 -- sleep 3600

kubectl set resources deployment web -n quota-demo \

  --requests=cpu=250m,memory=64Mi --limits=memory=128Mi

kubectl get deployment web -n quota-demo

kubectl describe replicaset -n quota-demo | grep -i -B1 -A2 "exceeded quota"

Only some of the six replicas become ready, because the namespace has 1 CPU of requests in total, which is already partly used by the first Pod. The ReplicaSet's events contain "exceeded quota" messages, even though the Deployment itself was accepted without any complaint.

  1. Free capacity by scaling down, and watch the missing Pods appear.

kubectl delete pod plain -n quota-demo

kubectl get pods -n quota-demo

kubectl describe resourcequota compute-quota -n quota-demo

  1. Clean up.

kubectl delete namespace quota-demo

11. Troubleshooting

Symptom or message

Likely cause and check

Pod rejected: must specify limits.memory or requests.cpu

A quota covers that resource and the Pod sets no value. Add resources to the Pod, or add a LimitRange with defaults

Pod rejected: exceeded quota ... requested, used, limited

The namespace has reached the hard limit. Check kubectl describe resourcequota, free capacity, or raise the quota through the proper process

Deployment created, but fewer Pods than expected

The ReplicaSet is blocked by quota or admission. Read the ReplicaSet events

Rollout stuck in the middle

No headroom for the surge Pods. Reduce maxSurge, or raise the quota temporarily

HPA wants more replicas than it gets

The quota blocks new Pods. Compare the HPA's maximum with the quota, and look at the ReplicaSet events

Pod rejected: maximum or minimum usage per Container

The value breaks the LimitRange. Change the Pod, or the LimitRange if the rule is wrong

Pod rejected: max limit to request ratio

The limit is too far above the request. Raise the request, or adjust maxLimitRequestRatio

Defaults not applied

The Pod was created before the LimitRange existed, the LimitRange is in another namespace, or a conflicting second LimitRange exists. Recreate the Pod and check the objects

Containers killed with OOMKilled after adding a LimitRange

The default memory limit is too low for the application. Set explicit values or raise the default

Quota shows usage but no Pods exist

Other objects count (such as claims, Services or Secrets), or usage lags briefly. Review all the resources in the quota status

Quota seems not enforced

No quota object exists in the namespace, the resource is not covered by the quota, or the ResourceQuota admission plugin is disabled on a self-managed cluster

Cannot create a PersistentVolumeClaim: exceeded quota

The storage or claim-count quota is full. Use the per-class quota keys to find which class is full

New quota does not take effect immediately

The quota controller needs a moment to compute usage. Check the status again after a short time

Pods stay Pending although quota has room

A different limit applies: the cluster itself has no free node capacity. Quota is a ceiling, not a reservation

 

12. Best Practices

  • Deploy a ResourceQuota and a LimitRange together in every shared namespace, and create them automatically with the namespace.

  • Use the LimitRange to supply sensible defaults and sanity limits, so that developers do not need to know about the quota to deploy.

  • Prefer quotas on requests for CPU, and be careful with quotas on CPU limits, because they force limits that can throttle applications.

  • Quota memory limits as well as requests, so that one namespace cannot over-commit node memory dangerously.

  • Cap object counts that stress the control plane or cost money, such as Services of type LoadBalancer, Secrets, ConfigMaps and claims.

  • Offer a few standard tiers, base them on measured usage, and review them regularly.

  • Leave headroom for rolling updates, autoscaling and Jobs, and keep HPA maximums consistent with the quota.

  • Prevent BestEffort Pods in production with a BestEffort-scoped quota, or with defaults that make every Pod Burstable at least.

  • Remember that quotas are ceilings, not reservations, and track real cluster capacity separately.

  • Alert on quota utilisation above about 80 or 90 percent, and on quota errors in ReplicaSet events.

  • Keep quota and LimitRange manifests in version control, with the change process for increases clearly documented.

  • Combine them with Pod Security labels, RBAC and network policies in the same namespace template.

13. Conclusion

ResourceQuotas and LimitRanges turn a shared Kubernetes cluster from a free-for-all into a managed environment. The LimitRange looks at each container, Pod and claim: it fills in defaults for the forgetful and rejects values that are unreasonably small or large. The ResourceQuota looks at the whole namespace: it caps total resources and counts of objects, so that no team can take more than its share. Both act at admission time, which makes them predictable, and both are best used together.

Design them as part of the namespace template, size them from real data, leave room for rollouts and autoscaling, monitor their utilisation and treat increases as a deliberate process. With these habits in place, capacity problems show up as early, readable signals, instead of as outages that affect other teams.