Kubernetes Resource Management and Scheduling
Kubernetes Resource Management and Scheduling
bitcodematrix.com | Kubernetes Series
1. Introduction
A Kubernetes cluster is a shared pool of CPU, memory, storage and special hardware. Two questions decide how well that pool is used: how much of it each workload is allowed to take (resource management), and which node each workload lands on (scheduling). Get them right and applications are stable, efficient and resilient to failures. Get them wrong and you see Pending Pods, noisy neighbours, evictions, uneven nodes and a large cloud bill.
This article ties the topics together. It explains how the scheduler decides where Pods go, how requests, limits, quotas and QoS shape that decision, how you steer placement with selectors, affinity, taints and topology spread, how priority and preemption work, and how autoscaling closes the loop. Earlier articles in this series cover requests and limits, QoS classes, OOMKilled, ResourceQuotas and LimitRanges in depth; here they are placed in the bigger picture.
After reading it you should be able to:
Describe the scheduler's filter, score and bind steps.
Explain node capacity versus allocatable and why Pods stay Pending.
Choose between nodeSelector, node affinity, Pod affinity, taints and topology spread.
Use priority classes and understand preemption.
Combine quotas, autoscalers and policies into a coherent resource strategy, and troubleshoot scheduling failures.
2. The Big Picture
Resource management and scheduling operate at three levels. Each level answers a different question and is controlled by different objects.
The glue between the levels is the request. The scheduler places Pods by comparing requests to node capacity, quotas cap the sum of requests and limits, autoscalers react to requests that cannot be satisfied or to utilisation measured against requests. Accurate requests are therefore the foundation of everything else in this article.
3. Node Capacity and Allocatable Resources
Every node reports its total capacity, but not all of it is available to Pods. The scheduler works with the allocatable amount:
Allocatable = Capacity - kube-reserved - system-reserved - eviction threshold
kube-reserved: resources set aside for Kubernetes daemons such as the kubelet and container runtime.
system-reserved: resources set aside for the operating system.
Eviction threshold: memory and disk that the kubelet tries to keep free.
kubectl describe node <node-name>
# Capacity:
# cpu: 4
# memory: 16Gi
# pods: 110
# Allocatable:
# cpu: 3920m
# memory: 15Gi
# pods: 110
# Allocated resources:
# cpu 2100m (53%) memory 6Gi (40%)
The numbers above are illustrative. Note the Pod count: each node has a maximum number of Pods (110 by default in the kubelet configuration, though it can be changed and some cloud providers set lower values), which can become the limiting factor on nodes with small Pods. Also remember that Allocated resources shows the sum of requests, not real usage.
4. How the Scheduler Works
The kube-scheduler is a control plane component that watches for Pods with no node assigned and picks a node for each. It does not start containers itself; it only writes the decision, and the kubelet on the chosen node does the rest. The control plane article in this series shows where the scheduler sits.
Figure 1: The scheduling cycle for one Pod
4.1 The steps
Queue: unscheduled Pods wait in a priority queue. Higher-priority Pods are taken first.
Filter: nodes that cannot run the Pod are removed. Typical checks are whether the requests fit, whether taints are tolerated, whether the nodeSelector and affinity rules match, whether required volumes can be attached in the right zone, and whether host ports are free.
Score: each remaining node receives a score from several plugins. By default the resource-fit plugin favours nodes with more free resources (a spreading strategy), while other plugins consider preferred affinity, image locality and topology spread. The node with the highest total wins, with ties broken randomly.
Bind: the scheduler writes the chosen node into the Pod, and the kubelet takes over.
If no node passes the filter step, the Pod stays Pending and the scheduler records the reasons as an event. The scheduler may then try preemption, described in section 8.
4.2 The scheduling framework and profiles
Internally the scheduler is a set of plugins attached to extension points. Built-in plugins include NodeResourcesFit, NodeAffinity, TaintToleration, PodTopologySpread, InterPodAffinity, VolumeBinding and ImageLocality. Cluster administrators can configure scheduling profiles to enable, disable or re-weight plugins, for example to switch the resource scoring strategy from spreading Pods out to packing them tightly. A Pod can select a profile or an entirely separate scheduler through the schedulerName field. On managed Kubernetes services, you may have limited access to these settings.
4.3 What the scheduler does not do
It does not consider real-time usage; it uses requests.
It does not move running Pods. Once placed, a Pod stays on its node until it terminates or is evicted. Tools such as the descheduler can rebalance by evicting Pods so they are rescheduled.
It does not guarantee even utilisation, only that the filters pass and the scores are favourable at the moment of placement.
5. Resource Requests, Limits and QoS in Scheduling
Only requests influence placement. A Pod with requests of 500m CPU and 1Gi memory fits on a node only if the node has at least that much allocatable capacity left after subtracting the requests of the Pods already there. Limits are enforced later by the kubelet and the kernel. The effective request of a Pod is the sum of its regular containers (adjusted for init containers using the documented rules), plus Pod overhead if a RuntimeClass defines one.
For details, see the articles on requests versus limits and on QoS classes.
6. Controlling Where Pods Run
By default the scheduler is free to use any node that fits. Several mechanisms let you constrain or influence that choice. They fall into two groups: those that attract Pods to nodes, and those that repel them.
6.1 nodeSelector and node affinity
apiVersion: v1
kind: Pod
metadata:
name: affinity-demo
spec:
affinity:
nodeAffinity:
requiredDuringSchedulingIgnoredDuringExecution:
nodeSelectorTerms:
- matchExpressions:
- key: disktype
operator: In
values: ["ssd"]
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 50
preference:
matchExpressions:
- key: topology.kubernetes.io/zone
operator: In
values: ["zone-a"]
containers:
- name: app
image: nginx:1.27
The required rule is a filter: a node that does not match is excluded. The preferred rule is a score: matching nodes get extra points but are not mandatory. The suffix IgnoredDuringExecution means that already running Pods are not moved if node labels change later.
6.2 Pod affinity and anti-affinity
spec:
affinity:
podAntiAffinity:
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 100
podAffinityTerm:
topologyKey: kubernetes.io/hostname
labelSelector:
matchLabels:
app: web
This asks the scheduler to prefer nodes that are not already running a Pod labelled app=web, spreading replicas across nodes. The topologyKey defines what "together" means: a node (kubernetes.io/hostname), a zone (topology.kubernetes.io/zone) and so on. Inter-Pod affinity is computationally expensive on very large clusters, so topology spread constraints are often the better tool for plain spreading.
6.3 Topology spread constraints
spec:
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: DoNotSchedule
labelSelector:
matchLabels:
app: web
maxSkew is the maximum allowed difference in matching Pod count between the most and least populated domains. DoNotSchedule makes the rule a hard filter, while ScheduleAnyway makes it a soft preference. This is the recommended way to keep replicas evenly spread across zones for high availability.
6.4 Taints and tolerations
A taint on a node repels Pods that do not have a matching toleration. Taints have an effect:
# Dedicate a node to GPU workloads
kubectl taint nodes gpu-node-1 dedicated=gpu:NoSchedule
# Pod toleration
spec:
tolerations:
- key: "dedicated"
operator: "Equal"
value: "gpu"
effect: "NoSchedule"
A toleration only permits a Pod to run on a tainted node; it does not attract it. To force GPU Pods onto GPU nodes, combine the toleration with node affinity or a nodeSelector. Kubernetes also adds taints automatically for conditions such as not-ready nodes and memory pressure.
7. Governing Resources with Quotas and Defaults
Requests and limits are set per container, but nothing forces developers to set them sensibly. Two namespace-scoped objects fill the gap, both covered in a dedicated article:
LimitRange supplies default requests and limits and enforces per-container minimum, maximum and ratio rules.
ResourceQuota caps total requests, limits, object counts and storage in a namespace, and can be scoped by PriorityClass or QoS-related scopes.
Together they turn the cluster into a fair multi-tenant system. A common pattern is one namespace per team, with a ResourceQuota sized to the team's budget and a LimitRange that makes sure every container has realistic defaults.
8. Priority and Preemption
Not all workloads are equally important. A PriorityClass assigns an integer priority to Pods. Priority affects three things: the order in which Pods leave the scheduling queue, whether a Pod can preempt others, and the kubelet's ranking when it evicts Pods under node pressure.
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
name: high-priority
value: 100000
preemptionPolicy: PreemptLowerPriority
globalDefault: false
description: "Production services"
---
apiVersion: v1
kind: Pod
metadata:
name: important
spec:
priorityClassName: high-priority
containers:
- name: app
image: nginx:1.27
8.1 How preemption works
A high-priority Pod cannot be scheduled because no node has enough free resources.
The scheduler looks for a node where evicting one or more lower-priority Pods would make room.
It marks the chosen node as nominated for the Pod and sends the victims a termination signal, honouring their graceful termination period.
Once the room is free, the high-priority Pod is scheduled. The nominated node is a hint, and the Pod could still land elsewhere if the situation changes.
Setting preemptionPolicy to Never keeps the Pod's place in the queue ahead of lower-priority Pods but stops it from evicting others, which suits batch work that is important but not urgent. Two built-in classes, system-cluster-critical and system-node-critical, exist for core components. Use high priorities sparingly, and restrict who may use them with quotas scoped by PriorityClass, otherwise everyone will claim to be critical.
9. Autoscaling: Matching Capacity to Demand
These tools depend on each other and on accurate requests. HPA calculates utilisation as a percentage of the request, so a request that is too high or too low produces misleading scaling. The node autoscaler reacts to Pods that are unschedulable because of their requests, so oversized requests directly cost money. Avoid running HPA and VPA on the same metric for the same workload, since they can fight each other. A common safe combination is VPA in recommendation mode to tune requests, with HPA to handle load.
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: web
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: web
minReplicas: 3
maxReplicas: 20
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
10. Special Hardware and Advanced Topics
Extended resources and device plugins: GPUs and other devices are advertised by device plugins as resources such as nvidia.com/gpu. They are requested in the limits section as whole units and cannot be overcommitted.
HugePages: can be requested as a separate resource, with the node pre-configured for them.
Dynamic Resource Allocation: a newer, more flexible API for describing and claiming devices. Its maturity depends on the Kubernetes version, so check the documentation.
CPU and Topology Managers: node-level features that can give Guaranteed Pods exclusive CPUs and align CPU, memory and device placement on NUMA nodes.
Pod scheduling gates: let a controller hold a Pod back from scheduling until a condition is met.
Custom schedulers: a second scheduler can be run and selected with schedulerName for specialised needs such as batch or gang scheduling.
Descheduler: a separate project that evicts Pods so they can be placed again, helping with rebalancing after nodes are added.
11. Hands-On Walkthrough
This exercise creates a Pending Pod, then fixes it through node labels, and shows spreading. Use a test cluster with at least two nodes.
Label one node and create a Pod that requires that label before it exists.
Observe the Pending status and read the scheduler event.
Add the label and watch the Pod get scheduled.
Create a Deployment with topology spread and see replicas distributed.
# 1. Pod that requires disktype=ssd
cat <<EOF | kubectl apply -f -
apiVersion: v1
kind: Pod
metadata:
name: needs-ssd
spec:
nodeSelector:
disktype: ssd
containers:
- name: app
image: nginx:1.27
resources:
requests:
cpu: 100m
memory: 64Mi
EOF
# 2. See why it is Pending
kubectl get pod needs-ssd
kubectl describe pod needs-ssd | grep -A4 Events
# 3. Label a node; the Pod should then schedule
kubectl get nodes
kubectl label node <node-name> disktype=ssd
kubectl get pod needs-ssd -o wide
Expected result: before labelling, the Events section shows a FailedScheduling message similar to "0/3 nodes are available: 3 node(s) didn't match Pod's node affinity/selector". After labelling, the Pod is bound to the labelled node within seconds. Exact wording varies by version.
# 4. Spread replicas across nodes
cat <<EOF | kubectl apply -f -
apiVersion: apps/v1
kind: Deployment
metadata:
name: spread-demo
spec:
replicas: 4
selector:
matchLabels: { app: spread-demo }
template:
metadata:
labels: { app: spread-demo }
spec:
topologySpreadConstraints:
- maxSkew: 1
topologyKey: kubernetes.io/hostname
whenUnsatisfiable: DoNotSchedule
labelSelector:
matchLabels: { app: spread-demo }
containers:
- name: app
image: nginx:1.27
resources:
requests: { cpu: 50m, memory: 32Mi }
EOF
kubectl get pods -l app=spread-demo -o wide
# Clean up
kubectl delete deployment spread-demo
kubectl delete pod needs-ssd
kubectl label node <node-name> disktype-
12. Troubleshooting Scheduling Problems
The first step is always the same: describe the Pod and read the Events. The scheduler lists, per reason, how many nodes were rejected.
kubectl describe pod <pod> | sed -n '/Events:/,$p'
kubectl get events --field-selector reason=FailedScheduling
Another frequent surprise is a cluster that seems to have enough capacity but still has Pending Pods. This usually means the free capacity is fragmented across many nodes (no single node can hold a large Pod), or that requests are much larger than real usage so the cluster is full on paper only. Compare the sum of requests (kubectl describe node) with actual usage (kubectl top nodes) to see which case you are in.
13. Best Practices
Set accurate CPU and memory requests for every container and review them regularly against real usage.
Use LimitRanges for defaults and ResourceQuotas for per-team budgets.
Spread replicas of important services across zones and nodes using topology spread constraints.
Prefer soft (preferred) rules unless a placement requirement is truly mandatory, to avoid unschedulable Pods.
Use taints and tolerations together with node affinity to dedicate nodes to special workloads.
Define a small, well-governed set of PriorityClasses and restrict high priorities.
Run Pod Disruption Budgets alongside spread rules so that voluntary disruptions do not remove too many replicas at once.
Combine HPA, VPA recommendations and a node autoscaler, and avoid conflicting scalers on the same metric.
Leave headroom for system components with kube-reserved and system-reserved settings.
Monitor Pending Pods, scheduling latency, node utilisation and the gap between requested and used resources.
Avoid nodeName and avoid very complex affinity rules that are hard to reason about and expensive to evaluate.
14. Conclusion
Resource management and scheduling are two halves of the same system. Requests describe what a workload needs; the scheduler uses them, together with selectors, affinity, taints, spread constraints and priority, to choose a node; limits, QoS and quotas keep the running workload within agreed boundaries; and autoscalers adjust the amount of capacity behind it all. When something goes wrong, the answer is nearly always in the Pod's events and in the comparison between requested and actual resources. Start with accurate requests, express placement needs as simply as possible, spread what must be highly available, and let governance and autoscaling handle growth. The other articles in this series on requests and limits, QoS, OOMKilled and quotas then give you the detail behind each building block.