kubernetes

Kubernetes HPA vs VPA vs Cluster Autoscaler

By Shubhankar Tripathi • • 5 min read

Kubernetes HPA vs VPA vs Cluster Autoscaler

bitcodematrix.com | Kubernetes Series

1. Introduction

Kubernetes offers three well-known autoscalers: the Horizontal Pod Autoscaler (HPA), the Vertical Pod Autoscaler (VPA) and the Cluster Autoscaler (CA). Their names sound alike, they all deal with "scaling", and newcomers often wonder which one to use. The honest answer is that they solve different problems at different layers, and a production platform usually uses more than one of them together.

This article compares the three side by side: what each one scales, what signal it reacts to, how it applies changes, where it is strong and where it is weak. It then shows how they cooperate during a traffic surge, which combinations are safe and which conflict, and how to choose for common workload types. Each autoscaler has its own detailed article in this series; this one is the map that connects them.

After reading it you should be able to:

  • State in one sentence what each autoscaler does and what it does not do.

  • Choose the right autoscaler (or combination) for a given workload.

  • Avoid the known conflicts, especially HPA with VPA on the same metric.

  • Trace what happens end to end when load increases and decreases.

  • Audit an existing cluster for autoscaling gaps.

2. The Three Autoscalers at a Glance

HPA scales Pod count, VPA scales Pod size, Cluster Autoscaler scales nodes, all linked through resource requests

Figure 1: What each autoscaler scales

Aspect

HPA

VPA

Cluster Autoscaler

Scales

Number of Pod replicas

Resource requests (and optionally limits) of containers

Number of nodes

Direction

Out and in (horizontal)

Up and down (vertical)

Out and in (infrastructure)

Reacts to

Metrics such as CPU, memory, custom or external metrics

Observed historical usage

Pending Pods that cannot be scheduled; underused nodes

Built into Kubernetes

Yes

No, separate add-on

No, separate add-on, often provided by cloud services

How changes apply

Changes replica count; no restart of existing Pods

Usually evicts and recreates Pods (restart); in-place where supported

Adds or deletes machines through the cloud provider

Typical reaction time

Seconds to a minute or two

Minutes to hours (needs history)

Minutes (node provisioning)

Needs

metrics-server (or adapter), CPU requests

metrics-server, history, Pod controller

Node groups, cloud permissions, accurate requests


3. Each Autoscaler in Brief

3.1 Horizontal Pod Autoscaler

The HPA keeps a workload's replica count in line with demand. It periodically compares a metric (for example average CPU utilisation) with a target and computes the desired number of replicas with the formula desired = ceil(current x currentValue / target). It then updates the scale of the Deployment, StatefulSet or other scalable target.

  • Strengths: fast, built in, no restarts, works with custom and external metrics, ideal for stateless services.

  • Limits: the application must tolerate several replicas; it cannot help when nodes are full; utilisation targets depend on accurate requests; scaling to zero needs extra tools.

3.2 Vertical Pod Autoscaler

The VPA right-sizes the resource requests of containers. Its Recommender learns from usage history, its Updater evicts Pods whose requests are far from the recommendation, and its Admission Controller sets the new values when Pods are recreated. It can also run in recommendation-only mode.

  • Strengths: removes guesswork from requests, reduces waste and OOM kills, useful for workloads that cannot scale out.

  • Limits: typically restarts Pods to apply changes; recommendations need time and history; cannot exceed node size; conflicts with HPA on the same metric; separate installation.

3.3 Cluster Autoscaler

The Cluster Autoscaler changes the number of nodes in node groups. It adds nodes when Pods are unschedulable because no node has enough free requests, and it removes nodes that have been underused and whose Pods can move elsewhere.

  • Strengths: converts Pod-level demand into infrastructure, controls cost by removing idle nodes, works with many cloud providers.

  • Limits: slow (machines must start), decisions based on requests not usage, scale-down blocked by many Pod settings, depends on node group design and cloud quotas.

4. How They Work Together

The autoscalers form a chain. Demand changes first affect Pods (HPA), the Pods need resources (sized by requests, tuned by the VPA), and the Pods need nodes (provided by the Cluster Autoscaler). The following timeline shows a typical traffic surge. The timings are illustrative, since real values depend on your configuration and cloud provider.

Phase

What happens

1. Surge begins

Traffic doubles; CPU utilisation of existing Pods rises above the HPA target.

2. HPA reacts (seconds to about a minute)

On its next evaluation the HPA computes a higher replica count and updates the Deployment; new Pods are created.

3. Pods wait

Some new Pods fit on existing nodes and start. Others are Pending with an Insufficient cpu or memory event.

4. Cluster Autoscaler reacts (next scan)

The CA sees Pending Pods, simulates a new node from a node group and requests an increase of that group.

5. Nodes arrive (typically minutes)

The cloud creates machines; they join and become Ready; the scheduler places the waiting Pods.

6. Load falls

The HPA lowers replicas, respecting its scale-down stabilisation window.

7. Nodes become idle

After the utilisation threshold, unneeded time and delays pass, the CA drains and deletes the underused nodes.

8. Background right-sizing

Over days, the VPA recommendations reveal that requests are too high or low; you or the VPA adjust them, which changes future HPA and CA behaviour.


Two practical lessons come out of this. First, the slowest link (node provisioning) determines how quickly you can really absorb a burst, so headroom matters. Second, requests influence every step, which is why the VPA's right-sizing improves the behaviour of the other two.

5. Which Combinations Are Safe?

Combination

Verdict

Notes

HPA + Cluster Autoscaler

Recommended

The standard pairing for stateless services; HPA adds Pods, CA adds nodes

VPA (Off mode) + HPA

Safe

Use VPA only for recommendations; apply them through your normal change process

VPA + Cluster Autoscaler

Works

VPA changes can change node needs; set maxAllowed to fit nodes

HPA on CPU + VPA on CPU (auto-applying)

Avoid

They fight: VPA changes the request that HPA uses as its denominator

HPA on custom metric + VPA on CPU and memory

Possible

Different signals; test carefully and watch restart frequency

All three together

Possible with care

Common in mature platforms: HPA on a business metric, VPA for requests, CA for nodes


The central rule is that two controllers should not act on the same signal for the same workload. Everything else is about timing and capacity.

6. Choosing by Workload Type

Workload

Suggested autoscaling

Reasoning

Stateless web or API service

HPA (CPU or request rate) + Cluster Autoscaler; VPA in Off mode for sizing

Scales out quickly; no restarts; right-sized requests improve HPA accuracy

Queue consumer or event processor

HPA on queue length (external metric) or an event-driven autoscaler such as KEDA, + Cluster Autoscaler

Demand is better measured by backlog than by CPU

Single-instance or hard-to-scale service (for example some databases)

VPA (recommendation or Initial mode) + sufficient node capacity

Cannot scale out; sizing matters; avoid surprise restarts

Batch jobs and CronJobs

VPA (Initial or Off) + Cluster Autoscaler

Each run starts fresh, so new requests apply without eviction; nodes scale with parallelism

Stateful clustered systems

Usually manual or operator-driven scaling; VPA for recommendations

Adding replicas needs data rebalancing; operators handle it better

GPU or special hardware workloads

Dedicated node groups with Cluster Autoscaler (or a provisioner); HPA on a relevant metric

Expensive nodes: scale-to-low is valuable, but start-up time is long

Development and test clusters

Cluster Autoscaler with aggressive scale-down; fixed small replica counts

Cost savings matter more than burst capacity


7. Related and Alternative Tools

  • KEDA: an event-driven autoscaler that builds on the HPA, adds many event sources and supports scaling to and from zero.

  • Karpenter and provider provisioners: launch right-sized nodes directly and consolidate them, as an alternative to node groups with the Cluster Autoscaler. Availability depends on your cloud.

  • In-place Pod resize: a Kubernetes feature that lets resources change without recreating the Pod, reducing the main drawback of vertical scaling. Its maturity depends on the version, and VPA support is evolving.

  • Managed autopilot-style modes: some cloud services size nodes and Pods for you. Check what they automate before adding your own autoscalers.

  • Pod Disruption Budgets, priorities and overprovisioning: supporting features that make autoscaling safe and responsive.

8. A Combined Reference Configuration

The following example sketches a typical setup for a stateless service: an HPA scaling on CPU, a VPA in recommendation-only mode, a Pod Disruption Budget, and a Deployment without a fixed replica count. The Cluster Autoscaler is configured at the cluster level, with an optional overprovisioning Deployment as described in its own article.

# Deployment: no spec.replicas, so the HPA stays in control

apiVersion: apps/v1

kind: Deployment

metadata:

  name: web

  namespace: shop

spec:

  selector:

    matchLabels:

      app: web

  template:

    metadata:

      labels:

        app: web

    spec:

      containers:

        - name: app

          image: registry.example.com/shop/web:1.0

          resources:

            requests:

              cpu: 250m

              memory: 256Mi

            limits:

              memory: 512Mi

          readinessProbe:

            httpGet:

              path: /healthz

              port: 8080

---

apiVersion: autoscaling/v2

kind: HorizontalPodAutoscaler

metadata:

  name: web

  namespace: shop

spec:

  scaleTargetRef:

    apiVersion: apps/v1

    kind: Deployment

    name: web

  minReplicas: 3

  maxReplicas: 20

  metrics:

    - type: Resource

      resource:

        name: cpu

        target:

          type: Utilization

          averageUtilization: 65

  behavior:

    scaleDown:

      stabilizationWindowSeconds: 300

---

apiVersion: autoscaling.k8s.io/v1

kind: VerticalPodAutoscaler

metadata:

  name: web

  namespace: shop

spec:

  targetRef:

    apiVersion: apps/v1

    kind: Deployment

    name: web

  updatePolicy:

    updateMode: "Off"

---

apiVersion: policy/v1

kind: PodDisruptionBudget

metadata:

  name: web

  namespace: shop

spec:

  minAvailable: 2

  selector:

    matchLabels:

      app: web


How the pieces fit: the HPA adds and removes replicas based on CPU; the VPA only reports recommendations that you review and apply to the requests in Git; the PDB keeps at least two replicas during node drains by the Cluster Autoscaler; and the readiness probe makes sure that new Pods receive traffic only when they are ready. The image name is a placeholder.

9. Hands-On: Audit Your Cluster

This exercise helps you see which autoscalers exist and whether they are set up sensibly. All commands are read-only.

  1. List HPAs and look at their targets and replica counts.

  2. List VPAs and look at their update modes and recommendations.

  3. Check whether a Cluster Autoscaler or another node provisioner is running.

  4. Look for Pending Pods and recent scale events.

  5. Look for workloads that have an HPA but no CPU requests.

# 1. HPAs

kubectl get hpa -A

 

# 2. VPAs (if the CRD is installed)

kubectl get vpa -A

kubectl get vpa -A -o custom-columns=NS:.metadata.namespace,NAME:.metadata.name,MODE:.spec.updatePolicy.updateMode

 

# 3. Node autoscaling

kubectl -n kube-system get deploy | grep -i -E "autoscaler|karpenter"

kubectl get nodes

 

# 4. Pending Pods and scaling events

kubectl get pods -A --field-selector=status.phase=Pending

kubectl get events -A --field-selector reason=TriggeredScaleUp

 

# 5. Containers without CPU requests in a namespace

kubectl get pods -n shop -o json | \

  jq -r '.items[] | .metadata.name as $p | .spec.containers[] | select(.resources.requests.cpu == null) | "\($p)/\(.name)"'


Interpretation: an HPA with TARGETS showing unknown usually means missing requests or metrics. A workload with both an auto-applying VPA and a CPU-based HPA is a conflict to resolve. Pending Pods with no scale-up events suggest that the node autoscaler is absent, at its maximum or blocked. On managed services, the autoscaler components may be hidden from you, in which case check your provider's console. The jq command requires jq to be installed.

10. Common Pitfalls

Pitfall

Consequence

Remedy

HPA and auto-applying VPA both on CPU

Oscillation: replicas and requests keep changing

Use different metrics or VPA in Off mode

No CPU requests on HPA targets

HPA cannot calculate utilisation; shows unknown

Set requests on all containers

Requests far above real usage

HPA scales late; CA adds unneeded nodes; high cost

Right-size using VPA recommendations

Requests far below real usage

HPA scales too early; throttling; evictions; OOM kills

Raise requests; check QoS

HPA maxReplicas larger than cluster capacity or quota

Pending Pods

Align with node group maximums and quotas

Fixed spec.replicas in GitOps manifests

Deployment tool resets the HPA's decisions

Remove the replicas field

Restrictive PDBs and bare Pods

Cluster Autoscaler cannot remove nodes

Adjust PDBs; use controllers

Slow node provisioning during bursts

Pods Pending for minutes

Overprovisioning headroom; smaller images; warm pools if offered

VPA recommendation larger than any node

Pods stay Pending

Set maxAllowed to fit node sizes

Autoscaling on memory

Memory often does not fall after load drops

Prefer CPU or request-based metrics for HPA


11. Best Practices

  • Start from accurate requests: they are the common currency of all three autoscalers.

  • Use the HPA for load-driven scaling of stateless workloads, with a metric that tracks real demand.

  • Use the VPA first in recommendation mode, and move to Initial or Auto only for workloads that tolerate restarts.

  • Never let two controllers act on the same metric of the same workload.

  • Make sure the cluster can grow: configure node autoscaling, node group limits and cloud quotas to match HPA maximums.

  • Protect availability with readiness probes, Pod Disruption Budgets and sensible priorities.

  • Plan headroom for bursts, because node provisioning is the slowest step.

  • Scale down gradually: use HPA stabilisation, CA delays, and review what blocks node removal.

  • Monitor Pending Pods, scaling events, restarts, and cost, and load test your scaling paths regularly.

  • Document which autoscaler owns which workload so that teams do not add conflicting ones.

12. Conclusion

The three autoscalers answer three different questions. The HPA asks how many Pods the current load needs. The VPA asks how large each Pod should be. The Cluster Autoscaler asks how many machines are needed to hold all the Pods. They are layers of one system, linked by resource requests, and each one is only as good as the layer beneath it. Use the HPA with the Cluster Autoscaler as the default pairing for stateless services, use the VPA to keep requests honest, keep their signals separate, and design for the slowest step. With those principles, autoscaling becomes a predictable and cost-effective part of your platform instead of a source of surprises.