Kubernetes HPA vs VPA vs Cluster Autoscaler
Kubernetes HPA vs VPA vs Cluster Autoscaler
bitcodematrix.com | Kubernetes Series
1. Introduction
Kubernetes offers three well-known autoscalers: the Horizontal Pod Autoscaler (HPA), the Vertical Pod Autoscaler (VPA) and the Cluster Autoscaler (CA). Their names sound alike, they all deal with "scaling", and newcomers often wonder which one to use. The honest answer is that they solve different problems at different layers, and a production platform usually uses more than one of them together.
This article compares the three side by side: what each one scales, what signal it reacts to, how it applies changes, where it is strong and where it is weak. It then shows how they cooperate during a traffic surge, which combinations are safe and which conflict, and how to choose for common workload types. Each autoscaler has its own detailed article in this series; this one is the map that connects them.
After reading it you should be able to:
State in one sentence what each autoscaler does and what it does not do.
Choose the right autoscaler (or combination) for a given workload.
Avoid the known conflicts, especially HPA with VPA on the same metric.
Trace what happens end to end when load increases and decreases.
Audit an existing cluster for autoscaling gaps.
2. The Three Autoscalers at a Glance
Figure 1: What each autoscaler scales
3. Each Autoscaler in Brief
3.1 Horizontal Pod Autoscaler
The HPA keeps a workload's replica count in line with demand. It periodically compares a metric (for example average CPU utilisation) with a target and computes the desired number of replicas with the formula desired = ceil(current x currentValue / target). It then updates the scale of the Deployment, StatefulSet or other scalable target.
Strengths: fast, built in, no restarts, works with custom and external metrics, ideal for stateless services.
Limits: the application must tolerate several replicas; it cannot help when nodes are full; utilisation targets depend on accurate requests; scaling to zero needs extra tools.
3.2 Vertical Pod Autoscaler
The VPA right-sizes the resource requests of containers. Its Recommender learns from usage history, its Updater evicts Pods whose requests are far from the recommendation, and its Admission Controller sets the new values when Pods are recreated. It can also run in recommendation-only mode.
Strengths: removes guesswork from requests, reduces waste and OOM kills, useful for workloads that cannot scale out.
Limits: typically restarts Pods to apply changes; recommendations need time and history; cannot exceed node size; conflicts with HPA on the same metric; separate installation.
3.3 Cluster Autoscaler
The Cluster Autoscaler changes the number of nodes in node groups. It adds nodes when Pods are unschedulable because no node has enough free requests, and it removes nodes that have been underused and whose Pods can move elsewhere.
Strengths: converts Pod-level demand into infrastructure, controls cost by removing idle nodes, works with many cloud providers.
Limits: slow (machines must start), decisions based on requests not usage, scale-down blocked by many Pod settings, depends on node group design and cloud quotas.
4. How They Work Together
The autoscalers form a chain. Demand changes first affect Pods (HPA), the Pods need resources (sized by requests, tuned by the VPA), and the Pods need nodes (provided by the Cluster Autoscaler). The following timeline shows a typical traffic surge. The timings are illustrative, since real values depend on your configuration and cloud provider.
Two practical lessons come out of this. First, the slowest link (node provisioning) determines how quickly you can really absorb a burst, so headroom matters. Second, requests influence every step, which is why the VPA's right-sizing improves the behaviour of the other two.
5. Which Combinations Are Safe?
The central rule is that two controllers should not act on the same signal for the same workload. Everything else is about timing and capacity.
6. Choosing by Workload Type
7. Related and Alternative Tools
KEDA: an event-driven autoscaler that builds on the HPA, adds many event sources and supports scaling to and from zero.
Karpenter and provider provisioners: launch right-sized nodes directly and consolidate them, as an alternative to node groups with the Cluster Autoscaler. Availability depends on your cloud.
In-place Pod resize: a Kubernetes feature that lets resources change without recreating the Pod, reducing the main drawback of vertical scaling. Its maturity depends on the version, and VPA support is evolving.
Managed autopilot-style modes: some cloud services size nodes and Pods for you. Check what they automate before adding your own autoscalers.
Pod Disruption Budgets, priorities and overprovisioning: supporting features that make autoscaling safe and responsive.
8. A Combined Reference Configuration
The following example sketches a typical setup for a stateless service: an HPA scaling on CPU, a VPA in recommendation-only mode, a Pod Disruption Budget, and a Deployment without a fixed replica count. The Cluster Autoscaler is configured at the cluster level, with an optional overprovisioning Deployment as described in its own article.
# Deployment: no spec.replicas, so the HPA stays in control
apiVersion: apps/v1
kind: Deployment
metadata:
name: web
namespace: shop
spec:
selector:
matchLabels:
app: web
template:
metadata:
labels:
app: web
spec:
containers:
- name: app
image: registry.example.com/shop/web:1.0
resources:
requests:
cpu: 250m
memory: 256Mi
limits:
memory: 512Mi
readinessProbe:
httpGet:
path: /healthz
port: 8080
---
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: web
namespace: shop
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: web
minReplicas: 3
maxReplicas: 20
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 65
behavior:
scaleDown:
stabilizationWindowSeconds: 300
---
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: web
namespace: shop
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: web
updatePolicy:
updateMode: "Off"
---
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: web
namespace: shop
spec:
minAvailable: 2
selector:
matchLabels:
app: web
How the pieces fit: the HPA adds and removes replicas based on CPU; the VPA only reports recommendations that you review and apply to the requests in Git; the PDB keeps at least two replicas during node drains by the Cluster Autoscaler; and the readiness probe makes sure that new Pods receive traffic only when they are ready. The image name is a placeholder.
9. Hands-On: Audit Your Cluster
This exercise helps you see which autoscalers exist and whether they are set up sensibly. All commands are read-only.
List HPAs and look at their targets and replica counts.
List VPAs and look at their update modes and recommendations.
Check whether a Cluster Autoscaler or another node provisioner is running.
Look for Pending Pods and recent scale events.
Look for workloads that have an HPA but no CPU requests.
# 1. HPAs
kubectl get hpa -A
# 2. VPAs (if the CRD is installed)
kubectl get vpa -A
kubectl get vpa -A -o custom-columns=NS:.metadata.namespace,NAME:.metadata.name,MODE:.spec.updatePolicy.updateMode
# 3. Node autoscaling
kubectl -n kube-system get deploy | grep -i -E "autoscaler|karpenter"
kubectl get nodes
# 4. Pending Pods and scaling events
kubectl get pods -A --field-selector=status.phase=Pending
kubectl get events -A --field-selector reason=TriggeredScaleUp
# 5. Containers without CPU requests in a namespace
kubectl get pods -n shop -o json | \
jq -r '.items[] | .metadata.name as $p | .spec.containers[] | select(.resources.requests.cpu == null) | "\($p)/\(.name)"'
Interpretation: an HPA with TARGETS showing unknown usually means missing requests or metrics. A workload with both an auto-applying VPA and a CPU-based HPA is a conflict to resolve. Pending Pods with no scale-up events suggest that the node autoscaler is absent, at its maximum or blocked. On managed services, the autoscaler components may be hidden from you, in which case check your provider's console. The jq command requires jq to be installed.
10. Common Pitfalls
11. Best Practices
Start from accurate requests: they are the common currency of all three autoscalers.
Use the HPA for load-driven scaling of stateless workloads, with a metric that tracks real demand.
Use the VPA first in recommendation mode, and move to Initial or Auto only for workloads that tolerate restarts.
Never let two controllers act on the same metric of the same workload.
Make sure the cluster can grow: configure node autoscaling, node group limits and cloud quotas to match HPA maximums.
Protect availability with readiness probes, Pod Disruption Budgets and sensible priorities.
Plan headroom for bursts, because node provisioning is the slowest step.
Scale down gradually: use HPA stabilisation, CA delays, and review what blocks node removal.
Monitor Pending Pods, scaling events, restarts, and cost, and load test your scaling paths regularly.
Document which autoscaler owns which workload so that teams do not add conflicting ones.
12. Conclusion
The three autoscalers answer three different questions. The HPA asks how many Pods the current load needs. The VPA asks how large each Pod should be. The Cluster Autoscaler asks how many machines are needed to hold all the Pods. They are layers of one system, linked by resource requests, and each one is only as good as the layer beneath it. Use the HPA with the Cluster Autoscaler as the default pairing for stateless services, use the VPA to keep requests honest, keep their signals separate, and design for the slowest step. With those principles, autoscaling becomes a predictable and cost-effective part of your platform instead of a source of surprises.