Kubernetes Horizontal Pod Autoscaler Explained
Kubernetes Horizontal Pod Autoscaler Explained
bitcodematrix.com | Kubernetes Series
1. Introduction
Traffic is rarely constant. A web service may be quiet at night and busy at lunchtime, a queue consumer may face sudden bursts, and a campaign can multiply load within minutes. Running enough replicas for the peak all day wastes money, while running too few causes slow responses and errors. The Horizontal Pod Autoscaler (HPA) solves this by automatically adjusting the number of Pod replicas of a workload based on observed metrics.
This article explains how the HPA works, how it calculates the number of replicas, how to configure it with the autoscaling/v2 API, how to tune scaling speed with the behavior field, how it interacts with requests, the Vertical Pod Autoscaler, the Cluster Autoscaler and Pod Disruption Budgets, and how to troubleshoot it. It builds on the earlier articles on requests and limits and on resource management and scheduling.
After reading it you should be able to:
Explain the HPA control loop and its scaling formula.
Write an HPA using resource, custom and external metrics.
Control scale-up and scale-down speed to avoid flapping.
Combine HPA with the other autoscalers safely.
Diagnose why an HPA shows unknown metrics or does not scale.
2. Horizontal vs Vertical Scaling
There are two ways to give an application more capacity. Horizontal scaling adds more Pods; vertical scaling gives each Pod more CPU or memory.
The HPA is built into Kubernetes as a control loop in the controller manager and a HorizontalPodAutoscaler API object. The VPA is a separate add-on project.
3. How the HPA Works
Figure 1: The HPA control loop
The HPA controller wakes up periodically (every 15 seconds by default) and reads the HorizontalPodAutoscaler objects.
For each HPA it queries the metrics API for the Pods selected by the target workload.
It computes the desired number of replicas from the current and target metric values.
It applies limits (minReplicas, maxReplicas, and any scaling behavior policies) and stabilisation rules.
It updates the scale subresource of the target (for example a Deployment), and the workload controller creates or deletes Pods.
3.1 The scaling formula
The core calculation, as documented by Kubernetes, is:
desiredReplicas = ceil[ currentReplicas x ( currentMetricValue / desiredMetricValue ) ]
Worked examples with a CPU utilisation target of 60 percent:
To avoid constant small changes, the controller ignores differences within a tolerance (0.1, or 10 percent, by default). In the last row the ratio 1.05 is within tolerance, so no scaling happens. The tolerance is a controller-manager setting that you can usually change only on self-managed clusters.
3.2 Utilisation is relative to requests
For resource metrics, the Utilization target is calculated as a percentage of the Pod's resource request, averaged over the Pods. A Pod that requests 200m CPU and uses 100m is at 50 percent. Two consequences follow. First, containers must declare requests, or the HPA cannot calculate utilisation. Second, a request that is too high makes the HPA scale late, while one that is too low makes it scale early and heavily. Accurate requests are essential, as explained in the article on requests and limits.
3.3 Pods that are not ready
The HPA treats Pods carefully while they start. Pods that are still starting up or that do not yet have metrics are handled conservatively so that a brief CPU spike during start-up does not trigger additional scaling. Several controller-manager settings control the initial readiness window. Readiness probes therefore matter: a Pod that reports Ready too early may be counted before it can serve traffic.
4. Prerequisites
The workload must be scalable: Deployment, StatefulSet, ReplicaSet or any resource with a scale subresource. DaemonSets cannot be scaled by the HPA because their replica count follows the nodes.
A metrics source: metrics-server for CPU and memory; a metrics adapter (for example the Prometheus Adapter) for custom metrics; an external metrics provider (for example KEDA) for metrics from outside the cluster.
Resource requests on every container whose utilisation is being measured.
Enough capacity: the scheduler needs room for new Pods, or a node autoscaler needs to add nodes.
# Verify that the metrics API works
kubectl top nodes
kubectl top pods -n shop
# Check the metrics API registration
kubectl get apiservices | grep metrics
5. Defining an HPA
5.1 The quick way
kubectl autoscale deployment web --cpu-percent=70 --min=3 --max=10
kubectl get hpa
5.2 The declarative way (autoscaling/v2)
apiVersion: autoscaling/v2
kind: HorizontalPodAutoscaler
metadata:
name: web
namespace: shop
spec:
scaleTargetRef:
apiVersion: apps/v1
kind: Deployment
name: web
minReplicas: 3
maxReplicas: 10
metrics:
- type: Resource
resource:
name: cpu
target:
type: Utilization
averageUtilization: 70
- type: Resource
resource:
name: memory
target:
type: Utilization
averageUtilization: 80
When several metrics are listed, the HPA calculates a desired replica count for each and uses the highest. This ensures that every metric stays within its target. The older autoscaling/v1 API supports only CPU utilisation, so use autoscaling/v2 for anything else.
5.3 Metric types
Target types are Utilization (a percentage of requests, resource metrics only), AverageValue (a value per Pod) and Value (a raw value for the whole object or external metric). The ContainerResource type is useful in Pods with sidecars, because a busy or idle sidecar would otherwise distort the Pod-level average.
metrics:
- type: Pods
pods:
metric:
name: httprequestspersecond
target:
type: AverageValue
averageValue: "100"
- type: External
external:
metric:
name: queuemessages_ready
selector:
matchLabels:
queue: orders
target:
type: AverageValue
averageValue: "30"
The metric names above are examples. They exist only if your adapter exposes them, and the exact names depend on your monitoring setup.
6. Controlling Scaling Speed: the behavior Field
Without tuning, an HPA can react too quickly to a short spike or remove Pods too eagerly when load dips, a problem known as flapping. The behavior field lets you set separate rules for scaling up and scaling down.
spec:
behavior:
scaleUp:
stabilizationWindowSeconds: 0
policies:
- type: Percent
value: 100
periodSeconds: 60
- type: Pods
value: 4
periodSeconds: 60
selectPolicy: Max
scaleDown:
stabilizationWindowSeconds: 300
policies:
- type: Percent
value: 10
periodSeconds: 60
The defaults, as documented, allow scale-up quickly (up to doubling or adding four Pods per 15 seconds, whichever is larger) and scale-down gradually with a five-minute stabilisation window. A common tuning is a slow scale-down (for example 10 percent per minute) to protect against traffic returning right after a dip, and a fast scale-up for bursty workloads.
7. Interaction with Other Components
7.1 Do not fight the HPA over replicas
If a Deployment manifest in Git contains a fixed spec.replicas value, every apply (for example from a GitOps tool) resets the replica count and fights with the HPA. When using an HPA, remove the replicas field from the manifest, or configure your deployment tooling to ignore it.
7.2 Scaling to zero
The HPA requires minReplicas of at least 1 in normal operation. Scaling to zero is possible only through a feature gate (HPAScaleToZero) with object or external metrics, and its availability depends on the Kubernetes version. Tools such as KEDA are commonly used for scale-to-zero designs.
8. Hands-On Walkthrough
This exercise follows the well-known example from the Kubernetes documentation. You need a test cluster with metrics-server installed.
Deploy a CPU-intensive sample web server with a CPU request.
Create an HPA targeting 50 percent CPU.
Generate load and watch the replicas increase.
Stop the load and watch the replicas decrease after the stabilisation window.
cat <<EOF | kubectl apply -f -
apiVersion: apps/v1
kind: Deployment
metadata:
name: php-apache
spec:
selector:
matchLabels:
run: php-apache
template:
metadata:
labels:
run: php-apache
spec:
containers:
- name: php-apache
image: registry.k8s.io/hpa-example
ports:
- containerPort: 80
resources:
requests:
cpu: 200m
limits:
cpu: 500m
---
apiVersion: v1
kind: Service
metadata:
name: php-apache
spec:
selector:
run: php-apache
ports:
- port: 80
EOF
kubectl autoscale deployment php-apache --cpu-percent=50 --min=1 --max=10
kubectl get hpa php-apache -w
In a second terminal, start a load generator:
kubectl run -i --tty load-generator --rm --image=busybox:1.36 --restart=Never -- \
/bin/sh -c "while sleep 0.01; do wget -q -O- http://php-apache; done"
Expected result (illustrative): within a minute or two the TARGETS column rises above 50 percent, REPLICAS grows step by step toward the maximum, and the Deployment shows more Pods. Stop the load generator with Ctrl+C. After the scale-down stabilisation window (five minutes by default) the replica count falls back toward the minimum. Exact timings and numbers depend on your cluster.
kubectl get hpa php-apache
# NAME REFERENCE TARGETS MINPODS MAXPODS REPLICAS
# php-apache Deployment/php-apache 250%/50% 1 10 5
kubectl describe hpa php-apache
# Clean up
kubectl delete hpa,service,deployment php-apache
9. Reading HPA Status
The describe output is the best diagnostic tool. Look at the Metrics, Conditions and Events sections.
10. Troubleshooting
11. Best Practices
Always set accurate CPU requests; the HPA's Utilization target depends on them.
Choose a metric that actually tracks load: CPU for compute-bound services, requests per second or queue depth for others. Memory is rarely a good scaling signal.
Set minReplicas high enough for availability (at least two or three for production) and maxReplicas as a cost and safety cap.
Leave headroom in the target, for example 60 to 70 percent utilisation, so that new Pods have time to start before the existing ones saturate.
Use the behavior field to scale up quickly and down slowly.
Use correct readiness and startup probes so that new Pods only receive traffic when ready.
Make sure the cluster can add capacity: configure a node autoscaler and check quotas.
Do not combine HPA and VPA on the same resource metric.
Remove fixed replica counts from manifests managed by GitOps.
Load test to verify that scaling is fast enough for your real traffic patterns, and monitor HPA metrics, replica counts and Pending Pods.
Make applications tolerate frequent scaling: graceful shutdown, connection draining and idempotent work.
12. Conclusion
The Horizontal Pod Autoscaler is a simple idea implemented as a careful control loop: measure, compare with a target, compute the replica count, apply limits and repeat. Its quality depends on the inputs around it: accurate requests, a metric that reflects real load, readiness probes that tell the truth, and a cluster that can supply nodes quickly. Configure it with the autoscaling/v2 API, tune its speed with the behavior field, avoid conflicts with fixed replica counts and the VPA, and use describe and events to understand its decisions. Together with the Cluster Autoscaler, it lets an application follow its demand automatically, which is one of the most valuable operational features of Kubernetes.