Kubernetes Cluster Autoscaler Explained
Kubernetes Cluster Autoscaler Explained
bitcodematrix.com | Kubernetes Series
1. Introduction
The Horizontal Pod Autoscaler can add Pods in seconds, but Pods need somewhere to run. When every node is full, new Pods stay Pending no matter how many the HPA asks for. The opposite problem also exists: when load drops and Pods are removed, the nodes remain and you keep paying for idle machines. The Cluster Autoscaler (CA) closes this gap by adding nodes when Pods cannot be scheduled and removing nodes that are no longer needed.
This article explains how the Cluster Autoscaler decides to scale up and down, what its main settings are, which Pods prevent a node from being removed, how it relates to node groups and cloud providers, how it differs from newer node provisioners such as Karpenter, and how to troubleshoot it. It builds on the articles on resource management and scheduling, the HPA and the VPA.
After reading it you should be able to:
Explain the scale-up and scale-down logic of the Cluster Autoscaler.
Describe node groups, expanders and the key tuning flags.
Identify what blocks scale-down and how to allow or prevent it deliberately.
Design a cluster that scales smoothly with HPA, including spare headroom.
Diagnose why nodes were or were not added or removed.
2. Where the Cluster Autoscaler Fits
Kubernetes has autoscaling at three levels. They work together, and each depends on the one before it.
An important point is that the Cluster Autoscaler does not look at CPU or memory usage graphs. It reacts to scheduling: it adds nodes because Pods cannot be placed, and it removes nodes because their Pods' requests are low and those Pods can be placed elsewhere. This is why accurate requests matter so much, as covered in the article on requests and limits.
The Cluster Autoscaler is a separate project (in the kubernetes/autoscaler repository) that runs as a Deployment in the cluster, usually in kube-system. It talks to a cloud provider integration to change the size of node groups. On many managed Kubernetes services it is offered as a built-in or one-click option, in which case the provider runs and configures it for you.
3. Key Concepts
4. How Scale-Up Works
Figure 1: Cluster Autoscaler scale-up and scale-down flow
The CA scans the cluster regularly (every 10 seconds by default) for Pods that are Pending and unschedulable.
For each such Pod it simulates adding a node from each node group, using the template node, and checks whether the Pod would fit, including node selectors, affinity, taints and tolerations.
If one or more groups would help, the expander picks one, and the CA asks the cloud provider to increase that group's size by the number of nodes needed.
The cloud provider creates the machines. When they join the cluster and become Ready, the scheduler places the waiting Pods on them.
Some details are worth remembering. The CA only reacts to Pods that are unschedulable because of resources or constraints that a new node could solve. If a Pod is Pending because of a missing PersistentVolume, an impossible nodeSelector that no node group can satisfy, or a quota limit, adding nodes will not help and the CA will not scale up. Scale-up also takes real time: provisioning a virtual machine, joining the cluster, pulling images and passing readiness often takes several minutes in total.
4.1 Expanders
Expanders can be combined in a comma-separated list, and later ones break ties for earlier ones. For example, a configuration that prefers a priority order and then least waste is common in cost-optimised clusters.
5. How Scale-Down Works
Every scan, the CA also looks for nodes that might be removed.
A node is a candidate when its utilisation (sum of requests divided by allocatable) is below the threshold, which defaults to 50 percent. DaemonSet Pods and mirror Pods are generally excluded from this calculation.
The CA simulates whether all the removable Pods on the node could be rescheduled on other nodes. If not, the node is kept.
If the node remains unneeded for a continuous period (10 minutes by default), the CA taints it, drains it using the eviction API, and asks the cloud provider to delete it.
After a scale-up, scale-down evaluation pauses for a delay (10 minutes by default) to avoid removing nodes that were just added.
5.1 What prevents a node from being removed
# Allow the CA to evict a Pod that would otherwise block scale-down
metadata:
annotations:
cluster-autoscaler.kubernetes.io/safe-to-evict: "true"
# Protect a Pod from being evicted by the CA
metadata:
annotations:
cluster-autoscaler.kubernetes.io/safe-to-evict: "false"
6. Configuration
The Cluster Autoscaler is configured with command-line flags on its Deployment. The exact flags and their availability vary by version and cloud provider, so always consult the documentation for the version that matches your Kubernetes minor version. The following are among the most commonly used.
# Fragment of the cluster-autoscaler container spec (illustrative)
command:
- ./cluster-autoscaler
- --cloud-provider=aws
- --node-group-auto-discovery=asg:tag=k8s.io/cluster-autoscaler/enabled,k8s.io/cluster-autoscaler/my-cluster
- --expander=least-waste
- --balance-similar-node-groups
- --scale-down-utilization-threshold=0.5
- --scale-down-unneeded-time=10m
The example uses AWS tag-based discovery as an illustration. Other providers use different discovery mechanisms. The CA also needs IAM or equivalent permissions to read and modify the node groups, and Kubernetes RBAC permissions to list and evict Pods and update nodes.
7. Designing Clusters for the Cluster Autoscaler
7.1 Node groups
Keep nodes in a group identical. The CA assumes every node in a group has the same capacity and labels. Mixed instance types in one group lead to wrong simulations unless they are very similar in size.
Use separate groups for different needs: GPU nodes, spot nodes, large-memory nodes, and so on, with taints and labels to steer Pods.
One group per zone for zonal storage. A Pod that needs a volume in zone A must get a node in zone A. If a single multi-zone group adds a node in the wrong zone, the Pod stays Pending. Per-zone groups with --balance-similar-node-groups avoid this.
Set realistic minimum and maximum sizes. The maximum protects your budget and quotas, and the minimum provides baseline capacity.
7.2 Headroom with overprovisioning
Because scale-up takes minutes, a sudden burst can leave Pods Pending for a while. A well-known technique is to run low-priority placeholder Pods that reserve capacity. When real Pods arrive, they preempt the placeholders immediately, and the evicted placeholders then become Pending, which triggers the CA to add nodes in the background.
apiVersion: scheduling.k8s.io/v1
kind: PriorityClass
metadata:
name: overprovisioning
value: -1
globalDefault: false
description: "Placeholder Pods that reserve spare capacity"
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: overprovisioning
spec:
replicas: 2
selector:
matchLabels:
app: overprovisioning
template:
metadata:
labels:
app: overprovisioning
spec:
priorityClassName: overprovisioning
containers:
- name: pause
image: registry.k8s.io/pause:3.9
resources:
requests:
cpu: "1"
memory: 1Gi
The priority of the placeholders must be lower than that of normal Pods (here negative) and above the CA's cutoff for expendable Pods, which has a default value; otherwise they would not trigger scale-up. Check the documentation for your version. Size the placeholders to the amount of headroom you want to keep.
7.3 Work with the HPA
A typical flow during a traffic surge is: the HPA raises replicas, the new Pods that do not fit stay Pending, the CA adds nodes, and the Pods are scheduled. Total reaction time is the HPA delay plus node provisioning time. Keep HPA targets conservative (for example 60 to 70 percent), keep some headroom, and use readiness probes so that new Pods become useful quickly.
8. Cluster Autoscaler vs Karpenter and Provider Features
The Cluster Autoscaler is the long-standing, provider-neutral solution. Newer approaches, most notably Karpenter, take a different design: instead of resizing predefined node groups, they look at the Pending Pods and launch right-sized nodes directly through the cloud API, and they actively consolidate nodes to reduce cost. Some managed services also offer their own node auto-provisioning or fully managed node modes.
Which to choose depends on your cloud, your managed service and your team's experience. This article focuses on the Cluster Autoscaler, but the principles about requests, PDBs, priority and headroom apply to both.
9. Hands-On Walkthrough
This exercise uses placeholder Pods to trigger a scale-up. It needs a cluster with the Cluster Autoscaler enabled and at least one node group that can grow. Use a non-production cluster and be aware that it may create billable nodes.
Note the current number of nodes.
Create a Deployment of pause containers that each request more CPU than the free capacity.
Watch Pods become Pending, then watch new nodes appear.
Scale the Deployment to zero and watch the nodes being removed after the delays.
kubectl get nodes
cat <<EOF | kubectl apply -f -
apiVersion: apps/v1
kind: Deployment
metadata:
name: inflate
spec:
replicas: 0
selector:
matchLabels:
app: inflate
template:
metadata:
labels:
app: inflate
spec:
containers:
- name: pause
image: registry.k8s.io/pause:3.9
resources:
requests:
cpu: "1"
EOF
kubectl scale deployment inflate --replicas=10
kubectl get pods -l app=inflate
kubectl get nodes -w
Expected result (illustrative): some Pods are Pending with the event "0/N nodes are available: Insufficient cpu" followed by a TriggeredScaleUp event from the cluster-autoscaler. After a few minutes new nodes appear and the Pods are scheduled. Then scale back down:
kubectl scale deployment inflate --replicas=0
# After the unneeded time and delays (about 10 minutes or more):
kubectl get nodes
# Inspect the autoscaler status and events
kubectl -n kube-system describe configmap cluster-autoscaler-status
kubectl get events -A --field-selector reason=TriggeredScaleUp
kubectl delete deployment inflate
Expected result: the extra nodes are cordoned, drained and deleted after the scale-down delays. Timing depends on your configuration and cloud provider, and on managed services the status ConfigMap may not be visible to you.
10. Troubleshooting
Start with three sources: the Pending Pod's events, the autoscaler's status ConfigMap (when available) and the autoscaler's logs.
11. Best Practices
Set accurate resource requests on every Pod; the CA's decisions depend on them.
Match the Cluster Autoscaler version to your Kubernetes minor version.
Use identical nodes in each group and separate groups for distinct hardware, spot capacity and zones.
Define sensible minimum and maximum sizes and check cloud quotas in advance.
Run workloads under controllers, and avoid bare Pods and unnecessary local storage.
Set Pod Disruption Budgets that protect availability but still allow some disruption, so nodes can be drained.
Use priority classes and low-priority overprovisioning for burst headroom.
Choose an expander that matches your goal, such as least-waste for efficiency or priority for spot-first strategies.
Run the CA on nodes that it will not remove, for example a small dedicated group or the control plane on managed services, with high availability where supported.
Monitor Pending Pods, scale-up latency, node counts and the CA's own metrics and alert on failures.
Test scale-up and scale-down regularly with a controlled load, not first during a real incident.
12. Conclusion
The Cluster Autoscaler gives Kubernetes elastic infrastructure. It watches for Pods that cannot be scheduled, simulates new nodes from your node groups, adds the ones that help, and later removes nodes whose Pods can safely move elsewhere. Its decisions are based on requests, not on live usage, so correct requests, sensible node groups, realistic Pod Disruption Budgets and a little headroom are the ingredients of a smooth experience. Combined with the HPA for Pods and the VPA for right-sizing, it completes the autoscaling picture. Where your platform offers newer provisioners such as Karpenter, evaluate them too, but the same principles of accurate requests and well-designed disruption rules will apply.