kubernetes

Kubernetes Cluster Autoscaler Explained

By Shubhankar Tripathi • • 5 min read

Kubernetes Cluster Autoscaler Explained

bitcodematrix.com | Kubernetes Series

1. Introduction

The Horizontal Pod Autoscaler can add Pods in seconds, but Pods need somewhere to run. When every node is full, new Pods stay Pending no matter how many the HPA asks for. The opposite problem also exists: when load drops and Pods are removed, the nodes remain and you keep paying for idle machines. The Cluster Autoscaler (CA) closes this gap by adding nodes when Pods cannot be scheduled and removing nodes that are no longer needed.

This article explains how the Cluster Autoscaler decides to scale up and down, what its main settings are, which Pods prevent a node from being removed, how it relates to node groups and cloud providers, how it differs from newer node provisioners such as Karpenter, and how to troubleshoot it. It builds on the articles on resource management and scheduling, the HPA and the VPA.

After reading it you should be able to:

  • Explain the scale-up and scale-down logic of the Cluster Autoscaler.

  • Describe node groups, expanders and the key tuning flags.

  • Identify what blocks scale-down and how to allow or prevent it deliberately.

  • Design a cluster that scales smoothly with HPA, including spare headroom.

  • Diagnose why nodes were or were not added or removed.

2. Where the Cluster Autoscaler Fits

Kubernetes has autoscaling at three levels. They work together, and each depends on the one before it.

Autoscaler

Scales

Signal

Horizontal Pod Autoscaler

Number of Pods

Metrics such as CPU, requests per second or queue length

Vertical Pod Autoscaler

Resource requests per Pod

Observed historical usage

Cluster Autoscaler

Number of nodes

Pending Pods that cannot fit, and nodes that are underused


An important point is that the Cluster Autoscaler does not look at CPU or memory usage graphs. It reacts to scheduling: it adds nodes because Pods cannot be placed, and it removes nodes because their Pods' requests are low and those Pods can be placed elsewhere. This is why accurate requests matter so much, as covered in the article on requests and limits.

The Cluster Autoscaler is a separate project (in the kubernetes/autoscaler repository) that runs as a Deployment in the cluster, usually in kube-system. It talks to a cloud provider integration to change the size of node groups. On many managed Kubernetes services it is offered as a built-in or one-click option, in which case the provider runs and configures it for you.

3. Key Concepts

Term

Meaning

Node group

A set of identical nodes managed together, such as an AWS Auto Scaling group, a GCP managed instance group, or an Azure scale set. The CA scales node groups, not individual machines.

Min and max size

Limits you set per node group; the CA never goes outside them.

Template node

The CA's model of a new node in a group (CPU, memory, labels, taints), used to simulate whether a Pending Pod would fit.

Unschedulable Pod

A Pod the scheduler could not place, marked with the PodScheduled condition False and reason Unschedulable.

Expander

The strategy for choosing a node group when several could fit the Pod.

Utilisation

For scale-down: the sum of Pod requests on a node divided by its allocatable capacity (not real usage).


4. How Scale-Up Works

Scale-up from Pending Pods and scale-down from underused nodes

Figure 1: Cluster Autoscaler scale-up and scale-down flow

  1. The CA scans the cluster regularly (every 10 seconds by default) for Pods that are Pending and unschedulable.

  2. For each such Pod it simulates adding a node from each node group, using the template node, and checks whether the Pod would fit, including node selectors, affinity, taints and tolerations.

  3. If one or more groups would help, the expander picks one, and the CA asks the cloud provider to increase that group's size by the number of nodes needed.

  4. The cloud provider creates the machines. When they join the cluster and become Ready, the scheduler places the waiting Pods on them.

Some details are worth remembering. The CA only reacts to Pods that are unschedulable because of resources or constraints that a new node could solve. If a Pod is Pending because of a missing PersistentVolume, an impossible nodeSelector that no node group can satisfy, or a quota limit, adding nodes will not help and the CA will not scale up. Scale-up also takes real time: provisioning a virtual machine, joining the cluster, pulling images and passing readiness often takes several minutes in total.

4.1 Expanders

Expander

Behaviour

random

Choose a suitable node group at random (the default)

most-pods

Choose the group that lets the most Pods be scheduled

least-waste

Choose the group that leaves the least idle CPU and memory after scheduling

price

Choose the cheapest option (supported only for some cloud providers)

priority

Choose according to a priority list you provide in a ConfigMap, for example prefer spot groups, then on-demand

grpc

Delegate the choice to an external service


Expanders can be combined in a comma-separated list, and later ones break ties for earlier ones. For example, a configuration that prefers a priority order and then least waste is common in cost-optimised clusters.

5. How Scale-Down Works

Every scan, the CA also looks for nodes that might be removed.

  1. A node is a candidate when its utilisation (sum of requests divided by allocatable) is below the threshold, which defaults to 50 percent. DaemonSet Pods and mirror Pods are generally excluded from this calculation.

  2. The CA simulates whether all the removable Pods on the node could be rescheduled on other nodes. If not, the node is kept.

  3. If the node remains unneeded for a continuous period (10 minutes by default), the CA taints it, drains it using the eviction API, and asks the cloud provider to delete it.

  4. After a scale-up, scale-down evaluation pauses for a delay (10 minutes by default) to avoid removing nodes that were just added.

5.1 What prevents a node from being removed

Blocker

Detail and how to handle it

Pods without a controller

A bare Pod (not managed by a Deployment, StatefulSet, Job and so on) cannot be recreated elsewhere. Use controllers, or annotate the Pod as safe to evict.

Pods with local storage

Pods using emptyDir or hostPath may block removal unless annotated as safe to evict.

Restrictive Pod Disruption Budgets

If evicting a Pod would violate its PDB, the node stays. Review PDBs that allow zero disruptions.

Scheduling constraints

Pods whose nodeSelector, affinity, taints or topology rules cannot be met on other nodes block removal.

Annotation: safe-to-evict false

The Pod annotation cluster-autoscaler.kubernetes.io/safe-to-evict set to "false" explicitly protects the node.

Node annotation: scale-down-disabled

The annotation cluster-autoscaler.kubernetes.io/scale-down-disabled set to "true" on a node excludes it.

Pods in kube-system without a PDB

System Pods are by default not moved unless they have a PDB or are annotated safe to evict.

Minimum group size

The group is already at its configured minimum.


# Allow the CA to evict a Pod that would otherwise block scale-down

metadata:

  annotations:

    cluster-autoscaler.kubernetes.io/safe-to-evict: "true"

 

# Protect a Pod from being evicted by the CA

metadata:

  annotations:

    cluster-autoscaler.kubernetes.io/safe-to-evict: "false"


6. Configuration

The Cluster Autoscaler is configured with command-line flags on its Deployment. The exact flags and their availability vary by version and cloud provider, so always consult the documentation for the version that matches your Kubernetes minor version. The following are among the most commonly used.

Flag

Purpose

--cloud-provider

Which provider integration to use (for example aws, gce, azure)

--nodes=min:max:groupName

Define a node group and its size limits manually

--node-group-auto-discovery

Discover node groups by tags or labels instead of listing them

--expander

Strategy for choosing among node groups

--scan-interval

How often the CA evaluates the cluster (default 10 s)

--scale-down-utilization-threshold

Utilisation below which a node is a removal candidate (default 0.5)

--scale-down-unneeded-time

How long a node must be unneeded before removal (default 10 m)

--scale-down-delay-after-add

Pause after a scale-up before scale-down is considered (default 10 m)

--balance-similar-node-groups

Keep similar node groups (for example in different zones) at similar sizes

--skip-nodes-with-local-storage and --skip-nodes-with-system-pods

Control which nodes are skipped during scale-down

--max-node-provision-time

How long to wait for a new node before treating it as failed (default 15 m)


# Fragment of the cluster-autoscaler container spec (illustrative)

command:

  - ./cluster-autoscaler

  - --cloud-provider=aws

  - --node-group-auto-discovery=asg:tag=k8s.io/cluster-autoscaler/enabled,k8s.io/cluster-autoscaler/my-cluster

  - --expander=least-waste

  - --balance-similar-node-groups

  - --scale-down-utilization-threshold=0.5

  - --scale-down-unneeded-time=10m


The example uses AWS tag-based discovery as an illustration. Other providers use different discovery mechanisms. The CA also needs IAM or equivalent permissions to read and modify the node groups, and Kubernetes RBAC permissions to list and evict Pods and update nodes.

7. Designing Clusters for the Cluster Autoscaler

7.1 Node groups

  • Keep nodes in a group identical. The CA assumes every node in a group has the same capacity and labels. Mixed instance types in one group lead to wrong simulations unless they are very similar in size.

  • Use separate groups for different needs: GPU nodes, spot nodes, large-memory nodes, and so on, with taints and labels to steer Pods.

  • One group per zone for zonal storage. A Pod that needs a volume in zone A must get a node in zone A. If a single multi-zone group adds a node in the wrong zone, the Pod stays Pending. Per-zone groups with --balance-similar-node-groups avoid this.

  • Set realistic minimum and maximum sizes. The maximum protects your budget and quotas, and the minimum provides baseline capacity.

7.2 Headroom with overprovisioning

Because scale-up takes minutes, a sudden burst can leave Pods Pending for a while. A well-known technique is to run low-priority placeholder Pods that reserve capacity. When real Pods arrive, they preempt the placeholders immediately, and the evicted placeholders then become Pending, which triggers the CA to add nodes in the background.

apiVersion: scheduling.k8s.io/v1

kind: PriorityClass

metadata:

  name: overprovisioning

value: -1

globalDefault: false

description: "Placeholder Pods that reserve spare capacity"

---

apiVersion: apps/v1

kind: Deployment

metadata:

  name: overprovisioning

spec:

  replicas: 2

  selector:

    matchLabels:

      app: overprovisioning

  template:

    metadata:

      labels:

        app: overprovisioning

    spec:

      priorityClassName: overprovisioning

      containers:

        - name: pause

          image: registry.k8s.io/pause:3.9

          resources:

            requests:

              cpu: "1"

              memory: 1Gi


The priority of the placeholders must be lower than that of normal Pods (here negative) and above the CA's cutoff for expendable Pods, which has a default value; otherwise they would not trigger scale-up. Check the documentation for your version. Size the placeholders to the amount of headroom you want to keep.

7.3 Work with the HPA

A typical flow during a traffic surge is: the HPA raises replicas, the new Pods that do not fit stay Pending, the CA adds nodes, and the Pods are scheduled. Total reaction time is the HPA delay plus node provisioning time. Keep HPA targets conservative (for example 60 to 70 percent), keep some headroom, and use readiness probes so that new Pods become useful quickly.

8. Cluster Autoscaler vs Karpenter and Provider Features

The Cluster Autoscaler is the long-standing, provider-neutral solution. Newer approaches, most notably Karpenter, take a different design: instead of resizing predefined node groups, they look at the Pending Pods and launch right-sized nodes directly through the cloud API, and they actively consolidate nodes to reduce cost. Some managed services also offer their own node auto-provisioning or fully managed node modes.

Aspect

Cluster Autoscaler

Karpenter-style provisioners

Unit of scaling

Pre-defined node groups

Individual nodes chosen per workload

Instance choice

Fixed per group

Flexible; picks from many instance types

Cloud support

Many providers

Depends on the project and provider support; check current status

Scale-down

Utilisation threshold and drain

Consolidation and expiry policies

Maturity and familiarity

Very widely used

Rapidly adopted; features vary by provider


Which to choose depends on your cloud, your managed service and your team's experience. This article focuses on the Cluster Autoscaler, but the principles about requests, PDBs, priority and headroom apply to both.

9. Hands-On Walkthrough

This exercise uses placeholder Pods to trigger a scale-up. It needs a cluster with the Cluster Autoscaler enabled and at least one node group that can grow. Use a non-production cluster and be aware that it may create billable nodes.

  1. Note the current number of nodes.

  2. Create a Deployment of pause containers that each request more CPU than the free capacity.

  3. Watch Pods become Pending, then watch new nodes appear.

  4. Scale the Deployment to zero and watch the nodes being removed after the delays.

kubectl get nodes

 

cat <<EOF | kubectl apply -f -

apiVersion: apps/v1

kind: Deployment

metadata:

  name: inflate

spec:

  replicas: 0

  selector:

    matchLabels:

      app: inflate

  template:

    metadata:

      labels:

        app: inflate

    spec:

      containers:

        - name: pause

          image: registry.k8s.io/pause:3.9

          resources:

            requests:

              cpu: "1"

EOF

 

kubectl scale deployment inflate --replicas=10

kubectl get pods -l app=inflate

kubectl get nodes -w


Expected result (illustrative): some Pods are Pending with the event "0/N nodes are available: Insufficient cpu" followed by a TriggeredScaleUp event from the cluster-autoscaler. After a few minutes new nodes appear and the Pods are scheduled. Then scale back down:

kubectl scale deployment inflate --replicas=0

 

# After the unneeded time and delays (about 10 minutes or more):

kubectl get nodes

 

# Inspect the autoscaler status and events

kubectl -n kube-system describe configmap cluster-autoscaler-status

kubectl get events -A --field-selector reason=TriggeredScaleUp

 

kubectl delete deployment inflate


Expected result: the extra nodes are cordoned, drained and deleted after the scale-down delays. Timing depends on your configuration and cloud provider, and on managed services the status ConfigMap may not be visible to you.

10. Troubleshooting

Start with three sources: the Pending Pod's events, the autoscaler's status ConfigMap (when available) and the autoscaler's logs.

Symptom

Likely cause

What to check or do

Pod Pending, no scale-up

Pod would not fit even on a new node, or constraint no group satisfies

Check the NotTriggerScaleUp event; compare requests with node group instance size; check selectors and taints

Pod Pending, no scale-up, group at max

Node group has reached its maximum size

Raise the maximum or add another group; check cloud quotas

Scale-up requested but node never arrives

Cloud quota, capacity shortage, bad launch template or join failure

Check cloud provider events and instance state; node bootstrap logs

Pods Pending in one zone only

Zonal volume needs a node in a specific zone

Use per-zone node groups and balance-similar-node-groups

Nodes not removed when idle

Blockers: PDB, local storage, bare Pods, annotations, kube-system Pods

Check CA logs for the reason; fix the blocker or annotate Pods as safe to evict

Nodes removed too aggressively

Low threshold or short unneeded time; Pods with small requests

Tune thresholds; set PDBs; keep overprovisioning headroom

Utilisation looks low yet node is kept

CA uses requests, not usage; large requests keep the node above threshold

Right-size requests with VPA recommendations

Scale-up is slow

Provisioning time, image pull, delays

Use overprovisioning, pre-pulled images and smaller images

Cluster oscillates (add then remove)

Thresholds and delays too tight; spiky load

Increase scale-down delays; smooth HPA behavior

CA is not running or fails

Permissions, version mismatch or crash

Check Deployment, logs, IAM permissions and version compatibility with Kubernetes


11. Best Practices

  • Set accurate resource requests on every Pod; the CA's decisions depend on them.

  • Match the Cluster Autoscaler version to your Kubernetes minor version.

  • Use identical nodes in each group and separate groups for distinct hardware, spot capacity and zones.

  • Define sensible minimum and maximum sizes and check cloud quotas in advance.

  • Run workloads under controllers, and avoid bare Pods and unnecessary local storage.

  • Set Pod Disruption Budgets that protect availability but still allow some disruption, so nodes can be drained.

  • Use priority classes and low-priority overprovisioning for burst headroom.

  • Choose an expander that matches your goal, such as least-waste for efficiency or priority for spot-first strategies.

  • Run the CA on nodes that it will not remove, for example a small dedicated group or the control plane on managed services, with high availability where supported.

  • Monitor Pending Pods, scale-up latency, node counts and the CA's own metrics and alert on failures.

  • Test scale-up and scale-down regularly with a controlled load, not first during a real incident.

12. Conclusion

The Cluster Autoscaler gives Kubernetes elastic infrastructure. It watches for Pods that cannot be scheduled, simulates new nodes from your node groups, adds the ones that help, and later removes nodes whose Pods can safely move elsewhere. Its decisions are based on requests, not on live usage, so correct requests, sensible node groups, realistic Pod Disruption Budgets and a little headroom are the ingredients of a smooth experience. Combined with the HPA for Pods and the VPA for right-sizing, it completes the autoscaling picture. Where your platform offers newer provisioners such as Karpenter, evaluate them too, but the same principles of accurate requests and well-designed disruption rules will apply.