kubernetes

Kubernetes Horizontal Pod Autoscaler Explained

By Shubhankar Tripathi • • 5 min read

Kubernetes Horizontal Pod Autoscaler Explained

bitcodematrix.com | Kubernetes Series

1. Introduction

Traffic is rarely constant. A web service may be quiet at night and busy at lunchtime, a queue consumer may face sudden bursts, and a campaign can multiply load within minutes. Running enough replicas for the peak all day wastes money, while running too few causes slow responses and errors. The Horizontal Pod Autoscaler (HPA) solves this by automatically adjusting the number of Pod replicas of a workload based on observed metrics.

This article explains how the HPA works, how it calculates the number of replicas, how to configure it with the autoscaling/v2 API, how to tune scaling speed with the behavior field, how it interacts with requests, the Vertical Pod Autoscaler, the Cluster Autoscaler and Pod Disruption Budgets, and how to troubleshoot it. It builds on the earlier articles on requests and limits and on resource management and scheduling.

After reading it you should be able to:

  • Explain the HPA control loop and its scaling formula.

  • Write an HPA using resource, custom and external metrics.

  • Control scale-up and scale-down speed to avoid flapping.

  • Combine HPA with the other autoscalers safely.

  • Diagnose why an HPA shows unknown metrics or does not scale.

2. Horizontal vs Vertical Scaling

There are two ways to give an application more capacity. Horizontal scaling adds more Pods; vertical scaling gives each Pod more CPU or memory.

Aspect

Horizontal (HPA)

Vertical (VPA)

What changes

Number of replicas

Resource requests (and limits) of each Pod

Best for

Stateless services that can run many copies

Workloads that cannot be split, or for right-sizing requests

Disruption

None for existing Pods; new Pods are added or removed

Pods are typically restarted to apply new values (unless in-place resize is used)

Limit

Application must tolerate multiple replicas

Limited by the size of a single node


The HPA is built into Kubernetes as a control loop in the controller manager and a HorizontalPodAutoscaler API object. The VPA is a separate add-on project.

3. How the HPA Works

Pods report metrics, HPA computes replicas and scales the workload

Figure 1: The HPA control loop

  1. The HPA controller wakes up periodically (every 15 seconds by default) and reads the HorizontalPodAutoscaler objects.

  2. For each HPA it queries the metrics API for the Pods selected by the target workload.

  3. It computes the desired number of replicas from the current and target metric values.

  4. It applies limits (minReplicas, maxReplicas, and any scaling behavior policies) and stabilisation rules.

  5. It updates the scale subresource of the target (for example a Deployment), and the workload controller creates or deletes Pods.

3.1 The scaling formula

The core calculation, as documented by Kubernetes, is:

desiredReplicas = ceil[ currentReplicas x ( currentMetricValue / desiredMetricValue ) ]


Worked examples with a CPU utilisation target of 60 percent:

Current replicas

Current utilisation

Target

Desired replicas

4

90%

60%

ceil(4 x 1.5) = 6

6

20%

60%

ceil(6 x 0.333) = 2

5

63%

60%

ceil(5 x 1.05) = 6, but see tolerance below


To avoid constant small changes, the controller ignores differences within a tolerance (0.1, or 10 percent, by default). In the last row the ratio 1.05 is within tolerance, so no scaling happens. The tolerance is a controller-manager setting that you can usually change only on self-managed clusters.

3.2 Utilisation is relative to requests

For resource metrics, the Utilization target is calculated as a percentage of the Pod's resource request, averaged over the Pods. A Pod that requests 200m CPU and uses 100m is at 50 percent. Two consequences follow. First, containers must declare requests, or the HPA cannot calculate utilisation. Second, a request that is too high makes the HPA scale late, while one that is too low makes it scale early and heavily. Accurate requests are essential, as explained in the article on requests and limits.

3.3 Pods that are not ready

The HPA treats Pods carefully while they start. Pods that are still starting up or that do not yet have metrics are handled conservatively so that a brief CPU spike during start-up does not trigger additional scaling. Several controller-manager settings control the initial readiness window. Readiness probes therefore matter: a Pod that reports Ready too early may be counted before it can serve traffic.

4. Prerequisites

  • The workload must be scalable: Deployment, StatefulSet, ReplicaSet or any resource with a scale subresource. DaemonSets cannot be scaled by the HPA because their replica count follows the nodes.

  • A metrics source: metrics-server for CPU and memory; a metrics adapter (for example the Prometheus Adapter) for custom metrics; an external metrics provider (for example KEDA) for metrics from outside the cluster.

  • Resource requests on every container whose utilisation is being measured.

  • Enough capacity: the scheduler needs room for new Pods, or a node autoscaler needs to add nodes.

# Verify that the metrics API works

kubectl top nodes

kubectl top pods -n shop

 

# Check the metrics API registration

kubectl get apiservices | grep metrics


5. Defining an HPA

5.1 The quick way

kubectl autoscale deployment web --cpu-percent=70 --min=3 --max=10

kubectl get hpa


5.2 The declarative way (autoscaling/v2)

apiVersion: autoscaling/v2

kind: HorizontalPodAutoscaler

metadata:

  name: web

  namespace: shop

spec:

  scaleTargetRef:

    apiVersion: apps/v1

    kind: Deployment

    name: web

  minReplicas: 3

  maxReplicas: 10

  metrics:

    - type: Resource

      resource:

        name: cpu

        target:

          type: Utilization

          averageUtilization: 70

    - type: Resource

      resource:

        name: memory

        target:

          type: Utilization

          averageUtilization: 80


When several metrics are listed, the HPA calculates a desired replica count for each and uses the highest. This ensures that every metric stays within its target. The older autoscaling/v1 API supports only CPU utilisation, so use autoscaling/v2 for anything else.

5.3 Metric types

Type

Source

Example

Resource

CPU and memory of the Pods, from metrics-server

Average CPU at 70% of requests

ContainerResource

CPU or memory of one named container in the Pod

Scale on the app container, ignoring a sidecar

Pods

Custom per-Pod metric from a custom metrics adapter

Requests per second per Pod

Object

Metric describing another Kubernetes object

Requests per second on an Ingress

External

Metric from outside the cluster

Queue length in a message broker


Target types are Utilization (a percentage of requests, resource metrics only), AverageValue (a value per Pod) and Value (a raw value for the whole object or external metric). The ContainerResource type is useful in Pods with sidecars, because a busy or idle sidecar would otherwise distort the Pod-level average.

metrics:

  - type: Pods

    pods:

      metric:

        name: httprequestspersecond

      target:

        type: AverageValue

        averageValue: "100"

  - type: External

    external:

      metric:

        name: queuemessages_ready

        selector:

          matchLabels:

            queue: orders

      target:

        type: AverageValue

        averageValue: "30"


The metric names above are examples. They exist only if your adapter exposes them, and the exact names depend on your monitoring setup.

6. Controlling Scaling Speed: the behavior Field

Without tuning, an HPA can react too quickly to a short spike or remove Pods too eagerly when load dips, a problem known as flapping. The behavior field lets you set separate rules for scaling up and scaling down.

spec:

  behavior:

    scaleUp:

      stabilizationWindowSeconds: 0

      policies:

        - type: Percent

          value: 100

          periodSeconds: 60

        - type: Pods

          value: 4

          periodSeconds: 60

      selectPolicy: Max

    scaleDown:

      stabilizationWindowSeconds: 300

      policies:

        - type: Percent

          value: 10

          periodSeconds: 60


Field

Meaning

stabilizationWindowSeconds

The HPA looks back over this window and uses the most conservative recommendation: the highest recommended count when scaling down, the lowest when scaling up. Default is 300 seconds for scale-down and 0 for scale-up.

policies

Limits on how much change is allowed per period, as a number of Pods or a percentage of current replicas

selectPolicy

Max (default) picks the policy allowing the largest change, Min the smallest, Disabled turns scaling in that direction off

tolerance

Newer versions may allow a per-HPA tolerance; check your version's documentation


The defaults, as documented, allow scale-up quickly (up to doubling or adding four Pods per 15 seconds, whichever is larger) and scale-down gradually with a five-minute stabilisation window. A common tuning is a slow scale-down (for example 10 percent per minute) to protect against traffic returning right after a dip, and a fast scale-up for bursty workloads.

7. Interaction with Other Components

Component

How it interacts with the HPA

Resource requests

HPA utilisation is measured against requests; wrong requests mean wrong scaling

Cluster Autoscaler or Karpenter

HPA adds Pods; if no node has room, Pods become Pending and the node autoscaler adds nodes. Scale-out latency includes node start-up time.

Vertical Pod Autoscaler

Do not let HPA and VPA both act on CPU or memory of the same workload; use VPA in recommendation mode, or VPA on memory with HPA on a custom metric

Pod Disruption Budget

Protects availability during voluntary disruption such as node drains; it does not block HPA scale-down in the same way, so keep minReplicas high enough

Deployment rolling updates

HPA continues to work during rollouts; set maxSurge and maxUnavailable with the scaled replica count in mind

ResourceQuota

If the namespace quota is exhausted, new Pods cannot be created and the HPA will not achieve its target

KEDA

An event-driven autoscaler that creates and manages HPAs and can scale from or to zero for event sources


7.1 Do not fight the HPA over replicas

If a Deployment manifest in Git contains a fixed spec.replicas value, every apply (for example from a GitOps tool) resets the replica count and fights with the HPA. When using an HPA, remove the replicas field from the manifest, or configure your deployment tooling to ignore it.

7.2 Scaling to zero

The HPA requires minReplicas of at least 1 in normal operation. Scaling to zero is possible only through a feature gate (HPAScaleToZero) with object or external metrics, and its availability depends on the Kubernetes version. Tools such as KEDA are commonly used for scale-to-zero designs.

8. Hands-On Walkthrough

This exercise follows the well-known example from the Kubernetes documentation. You need a test cluster with metrics-server installed.

  1. Deploy a CPU-intensive sample web server with a CPU request.

  2. Create an HPA targeting 50 percent CPU.

  3. Generate load and watch the replicas increase.

  4. Stop the load and watch the replicas decrease after the stabilisation window.

cat <<EOF | kubectl apply -f -

apiVersion: apps/v1

kind: Deployment

metadata:

  name: php-apache

spec:

  selector:

    matchLabels:

      run: php-apache

  template:

    metadata:

      labels:

        run: php-apache

    spec:

      containers:

        - name: php-apache

          image: registry.k8s.io/hpa-example

          ports:

            - containerPort: 80

          resources:

            requests:

              cpu: 200m

            limits:

              cpu: 500m

---

apiVersion: v1

kind: Service

metadata:

  name: php-apache

spec:

  selector:

    run: php-apache

  ports:

    - port: 80

EOF

 

kubectl autoscale deployment php-apache --cpu-percent=50 --min=1 --max=10

kubectl get hpa php-apache -w


In a second terminal, start a load generator:

kubectl run -i --tty load-generator --rm --image=busybox:1.36 --restart=Never -- \

  /bin/sh -c "while sleep 0.01; do wget -q -O- http://php-apache; done"


Expected result (illustrative): within a minute or two the TARGETS column rises above 50 percent, REPLICAS grows step by step toward the maximum, and the Deployment shows more Pods. Stop the load generator with Ctrl+C. After the scale-down stabilisation window (five minutes by default) the replica count falls back toward the minimum. Exact timings and numbers depend on your cluster.

kubectl get hpa php-apache

# NAME         REFERENCE               TARGETS   MINPODS   MAXPODS   REPLICAS

# php-apache   Deployment/php-apache   250%/50%  1         10        5

 

kubectl describe hpa php-apache

 

# Clean up

kubectl delete hpa,service,deployment php-apache


9. Reading HPA Status

The describe output is the best diagnostic tool. Look at the Metrics, Conditions and Events sections.

Condition

Meaning

AbleToScale

Whether the HPA can fetch and update the scale of the target, and whether it is within a back-off or stabilisation period

ScalingActive

Whether metrics can be fetched and used for calculation; False with FailedGetResourceMetric means a metrics problem

ScalingLimited

Whether the desired count was capped by minReplicas, maxReplicas or a behavior policy


10. Troubleshooting

Symptom

Likely cause

What to check or do

TARGETS shows <unknown>/70%

Metrics not available: metrics-server missing or unhealthy, or Pods have no requests

kubectl top pods; check metrics-server logs; add resource requests

Event: missing request for cpu

A container has no CPU request

Add requests to all containers or use ContainerResource

HPA does not scale up although CPU is high

Already at maxReplicas, or ScalingLimited true, or within tolerance

kubectl describe hpa; raise maxReplicas; review policies

New Pods stay Pending

No node capacity or quota exhausted

Check scheduling events, node autoscaler and ResourceQuota

Scaling up and down repeatedly (flapping)

Spiky metric, short windows, slow Pod start-up

Increase scale-down stabilisation; use a smoother metric; fix readiness probes

Replicas reset after every deploy

Fixed spec.replicas in the manifest

Remove replicas from the manifest used with the HPA

Scaling is slower than expected

Default 15 s loop, start-up time, node provisioning, policies

Review behavior policies; pre-pull images; keep warm capacity

Custom metric not found

Adapter not installed or metric name wrong

kubectl get --raw for the custom metrics API path; check adapter configuration

Memory-based HPA does not reduce replicas

Memory does not fall when load falls in many runtimes

Prefer CPU or request-based metrics; memory is a poor scaling signal for most applications


11. Best Practices

  • Always set accurate CPU requests; the HPA's Utilization target depends on them.

  • Choose a metric that actually tracks load: CPU for compute-bound services, requests per second or queue depth for others. Memory is rarely a good scaling signal.

  • Set minReplicas high enough for availability (at least two or three for production) and maxReplicas as a cost and safety cap.

  • Leave headroom in the target, for example 60 to 70 percent utilisation, so that new Pods have time to start before the existing ones saturate.

  • Use the behavior field to scale up quickly and down slowly.

  • Use correct readiness and startup probes so that new Pods only receive traffic when ready.

  • Make sure the cluster can add capacity: configure a node autoscaler and check quotas.

  • Do not combine HPA and VPA on the same resource metric.

  • Remove fixed replica counts from manifests managed by GitOps.

  • Load test to verify that scaling is fast enough for your real traffic patterns, and monitor HPA metrics, replica counts and Pending Pods.

  • Make applications tolerate frequent scaling: graceful shutdown, connection draining and idempotent work.

12. Conclusion

The Horizontal Pod Autoscaler is a simple idea implemented as a careful control loop: measure, compare with a target, compute the replica count, apply limits and repeat. Its quality depends on the inputs around it: accurate requests, a metric that reflects real load, readiness probes that tell the truth, and a cluster that can supply nodes quickly. Configure it with the autoscaling/v2 API, tune its speed with the behavior field, avoid conflicts with fixed replica counts and the VPA, and use describe and events to understand its decisions. Together with the Cluster Autoscaler, it lets an application follow its demand automatically, which is one of the most valuable operational features of Kubernetes.