kubernetes

Kubernetes Autoscaling: How Applications Scale in Production

By Shubhankar Tripathi • • 5 min read

Kubernetes Autoscaling: How Applications Scale in Production

bitcodematrix.com | Kubernetes Series

1. Introduction

Enabling an autoscaler is the easy part of scaling in production. The hard part is making sure that when load arrives, new capacity is ready quickly enough, the application can use it safely, the databases and downstream services behind it can absorb the extra traffic, and the bill stays under control. Many "autoscaling" incidents are not caused by the autoscalers at all: the new Pods take four minutes to become ready, the database runs out of connections, the nodes cannot be provisioned because a cloud quota was reached, or the metric that drives scaling does not reflect real demand.

This article is the capstone of the autoscaling part of this series. The earlier articles explain the HPA, the VPA and the Cluster Autoscaler individually and compare them. Here we look at the whole system from the production point of view: the scale-out path and its time budget, choosing scaling signals, preparing applications and platform, protecting dependencies, controlling cost, observing and testing scaling, and handling failure modes.

After reading it you should be able to:

  • Describe every step between a traffic spike and new Pods serving requests, and estimate how long it takes.

  • Choose scaling signals and strategies (reactive, scheduled, event-driven) for different workloads.

  • Prepare applications (probes, graceful shutdown, pools) and the platform (headroom, quotas, images) for scaling.

  • Protect databases and downstream services from the load of a scale-out.

  • Monitor, load test and troubleshoot autoscaling, and apply a production readiness checklist.

2. Scaling Is a Chain, Not a Switch

For a request to be served by a new replica, a whole chain of things must work. Autoscalers handle only part of it.

Link in the chain

What must be true

Typical failure

Signal

A metric that reflects real demand is collected and available

Metric is stale, missing or does not track load

Pod autoscaling

HPA targets, limits and behavior are sensible

Already at maxReplicas; flapping; unknown metrics

Scheduling

Nodes have room, and constraints can be satisfied

Pending Pods due to requests, affinity, taints or zone volumes

Node capacity

Node autoscaler can add nodes quickly; cloud quota and IP addresses exist

Quota exhausted; node never joins; IP exhaustion

Start-up

Image is pulled and the application becomes ready fast

Large images; slow warm-up; failing probes

Traffic distribution

Load balancer and Service include new Pods promptly

Connections stay pinned to old Pods; slow endpoint propagation

Dependencies

Databases, caches and APIs can take the extra load

Connection limits; downstream rate limiting; cascading failure

Shutdown

Scale-in does not drop requests

Requests cut off when Pods terminate


The weakest link sets the real scaling speed and safety. The rest of this article works through the chain.

3. Scaling Strategies

There are three broad strategies, and mature platforms mix them.

Strategy

How it works

Best for

Reactive

Scale in response to measured load (HPA on CPU, request rate, queue depth)

Unpredictable or variable demand

Scheduled

Raise minimum capacity at known times, such as business hours or a planned sale, using scheduled scaling features of your platform or an event-driven autoscaler with a cron trigger

Predictable peaks that start faster than reactive scaling can follow

Predictive or manual pre-scaling

Capacity is increased ahead of a known event by people or tooling, based on forecasts or past events

Launches, campaigns, major events


Reactive scaling alone always lags demand by the time it takes to measure, decide, schedule and start. When a peak is known in advance, combine reactive scaling with a higher minimum during the peak window so that the reactive part only handles the surprise on top. An event-driven autoscaler such as KEDA can express this with a cron trigger. The following sketch is illustrative; check the documentation of the version you use.

apiVersion: keda.sh/v1alpha1

kind: ScaledObject

metadata:

  name: web-business-hours

  namespace: shop

spec:

  scaleTargetRef:

    name: web

  minReplicaCount: 3

  maxReplicaCount: 30

  triggers:

    - type: cron

      metadata:

        timezone: Asia/Kolkata

        start: "0 8 1-5"

        end: "0 20 1-5"

        desiredReplicas: "10"

    - type: cpu

      metricType: Utilization

      metadata:

        value: "65"


Here the workload keeps at least ten replicas during weekday business hours, and the CPU trigger can add more on top. Outside those hours it falls back to the minimum. Do not combine a KEDA ScaledObject with a separate HPA on the same workload; KEDA manages its own HPA.

4. The Scale-Out Path and Its Time Budget

Steps from load increase to new Pods serving traffic, with and without node provisioning

Figure 1: The scale-out path and where the time goes

Step

Typical duration (illustrative)

What influences it

Metrics collection and aggregation

Seconds to about a minute

Scrape interval, metrics-server or adapter delay

HPA decision

About 15 seconds per loop

Controller settings, behavior policies

Scheduling

Typically under a few seconds

Cluster size, constraints

Node provisioning (if needed)

Often one to several minutes

Cloud provider, instance type, node image and bootstrap

Image pull

Seconds to minutes

Image size, registry location, caching

Application start-up and warm-up

Seconds to minutes

Runtime, caches, JIT, migrations, probes


The numbers are examples, not promises. Measure your own path with a test (section 12). The practical conclusion is that you need to compare your measured time to new capacity with how fast your traffic can grow. If a spike can double traffic in 30 seconds and new capacity takes five minutes, you need headroom, scheduled capacity, or a lower target utilisation, not a faster HPA.

4.1 Ways to shorten the path

  • Lower the HPA target (for example 50 to 65 percent instead of 80) so that spare capacity exists while new Pods start.

  • Keep node headroom with low-priority overprovisioning Pods so that new Pods can be scheduled immediately.

  • Use small images and keep them in a registry close to the cluster; consider pre-pulling or image caching features of your platform.

  • Make the application start quickly: avoid heavy initialisation, defer warm-up, and use startup probes for slow starters.

  • Tune HPA behavior to allow fast scale-up (large step sizes, no stabilisation delay) and slower scale-down.

  • Right-size requests so that fewer Pods are needed and each fits easily.

5. Choosing Scaling Signals

The HPA is only as good as the metric it watches. A good signal rises when demand rises, falls when demand falls, is available quickly, and is proportional (doubling the load roughly doubles the metric per Pod).

Signal

Good for

Cautions

CPU utilisation

Compute-bound services

Requires accurate CPU requests; I/O-bound services may saturate before CPU rises

Requests per second per Pod

Web and API services

Needs a metrics adapter; set the target from load tests

Latency or error rate

Direct SLO protection

Can be noisy and may fall only after capacity is added; often better as an alert than as a scaling input

Concurrency or in-flight requests

Services with limited worker threads

Needs application metrics

Queue depth or lag

Workers and consumers

Scale on backlog per worker; use an external metric or an event-driven autoscaler

Memory

Rarely a good scaling signal

Often does not fall after load drops; use for VPA sizing instead

Business or custom metric

Domain-specific demand

Must be reliable and low-latency


When possible, derive the target from a load test: find the load per Pod at which latency starts to degrade, then set the target to roughly 60 to 70 percent of that. Treat these percentages as a starting point rather than a rule.

6. Preparing the Application

6.1 Stateless and horizontally safe

Horizontal scaling works only if replicas are interchangeable. Keep session state in a shared store or in signed tokens, avoid local files that other replicas need, and make background jobs safe to run on several replicas (leader election or locking where needed).

6.2 Probes that tell the truth

  • Readiness probe: returns success only when the Pod can serve real traffic. Without it, new Pods receive traffic too early and users see errors during scale-out.

  • Startup probe: protects slow-starting applications from being killed by liveness checks while they start.

  • Liveness probe: use carefully; an overly strict liveness probe under high load can cause restarts that make the load problem worse.

6.3 Graceful shutdown for scale-in

Scale-in and node drains terminate Pods. To avoid dropping requests, the Pod should stop receiving new traffic, finish current work, and exit. Kubernetes sends SIGTERM, and removal from Service endpoints happens in parallel, so there can be a short period in which traffic still arrives. A short preStop delay and a proper SIGTERM handler address this.

spec:

  terminationGracePeriodSeconds: 60

  containers:

    - name: app

      image: registry.example.com/shop/web:1.0

      lifecycle:

        preStop:

          exec:

            command: ["sleep", "10"]

      readinessProbe:

        httpGet:

          path: /healthz

          port: 8080

        periodSeconds: 5

      startupProbe:

        httpGet:

          path: /healthz

          port: 8080

        failureThreshold: 30

        periodSeconds: 2


The preStop sleep gives load balancers and endpoint controllers time to stop sending traffic before the application starts shutting down, and the application itself should drain connections when it receives SIGTERM. The sleep command requires the container image to include it. The image name and values are placeholders to adapt.

6.4 Requests, spreading and disruption budgets

  • Set realistic CPU and memory requests (see the article on requests and limits).

  • Use topology spread constraints so that replicas are spread across nodes and zones.

  • Use a Pod Disruption Budget so that voluntary disruptions, including node removal by the autoscaler, never take down too many replicas at once.

  • Choose a QoS class that matches the workload's importance.

7. Preparing the Platform

Area

What to check

Node groups

Separate groups for different hardware, spot versus on-demand, and zones; sensible min and max sizes; identical nodes in a group

Cloud quotas and limits

Instance quotas, load balancer limits, IP address and subnet capacity, per-node Pod limits; request increases before you need them

Namespace quotas

ResourceQuota values that allow the HPA maximum; LimitRange defaults

Headroom

Overprovisioning Pods or minimum node counts that match your burst needs

Image delivery

Small images, registry near the cluster, rate limits of public registries avoided with a mirror or cache

Cluster add-ons that must scale too

CoreDNS (consider a proportional autoscaler), ingress controllers, service mesh sidecars and gateways, metrics pipeline, log shippers

Metrics pipeline

metrics-server and any custom metrics adapter running with high availability and enough resources

Control plane

API server and etcd limits on large clusters; managed services publish their limits; heavy scaling creates many objects and events


A frequent hidden bottleneck is the add-ons. If the application scales to 200 replicas but CoreDNS, the ingress controller or the service mesh gateway stays at its original size, the add-on becomes the new limit.

8. Protecting Dependencies

Scaling out the application multiplies its demands on everything behind it. This is the most common production surprise: the web tier scales perfectly and takes the database down.

Dependency

Risk when you scale out

Mitigation

Relational database

Total connections exceed maxconnections

Cap maxReplicas using a connection budget; use a connection pooler; read replicas and caching

Cache

Cold new Pods cause cache stampedes; connection counts grow

Warm-up, request coalescing, sensible TTLs

Downstream APIs

Rate limits or quotas are hit

Client-side rate limiting, retries with back-off and jitter, circuit breakers

Message broker

More consumers than partitions give no benefit

Match consumer count to partitions or shards

Shared storage

Throughput or IOPS limits

Capacity planning; avoid shared volumes for hot paths


A simple connection budget calculation illustrates the point:

maxReplicas x connections per Pod  <=  database max connections - reserved connections

 

Example: 30 replicas x 10 connections = 300 connections

         database allows 400, reserve 100 for admin and jobs

         300 <= 300, so maxReplicas of 30 is the upper bound


Treat the HPA maxReplicas as a safety control derived from the capacity of the whole system, not just a cost cap. Add backpressure (queues, load shedding, timeouts) so that overload degrades gracefully instead of cascading.

9. Patterns by Workload Type

Workload

Pattern

Notes

Public web or API

HPA on CPU or request rate, scheduled minimums for known peaks, node autoscaling with headroom

Readiness probes and graceful shutdown are essential

Queue workers

Scale on queue depth or lag per worker, scale to a low minimum overnight

Make jobs idempotent; finish or return messages on shutdown

Batch and scheduled jobs

Cluster Autoscaler for nodes, VPA for requests, parallelism limits

Separate node group, often with spot capacity

Stateful services

Plan capacity manually or through operators; vertical sizing with care

Scaling needs data rebalancing and careful rollout

ML inference and GPU

Dedicated expensive nodes, scale on queue or latency, longer warm-up

Keep a small warm minimum if start-up is slow

Multi-tenant platform

Per-namespace quotas, priority classes and fair-share policies

Prevent one tenant from consuming all scaling headroom


10. Cost Control

Autoscaling reduces cost by removing idle capacity, but it can also hide waste or add cost if configured carelessly.

  • Right-size requests. Inflated requests make the HPA add Pods early and make nodes look full. Use VPA recommendations or usage data.

  • Set maxReplicas and node group maximums as budget and safety limits, and alert when they are reached.

  • Choose minimums deliberately. A high minimum is a fixed cost; a low one is a risk. Review them against real traffic.

  • Use cheaper capacity where it is safe, such as spot or preemptible nodes for fault-tolerant workloads, with PDBs and multiple instance types.

  • Allow scale-down. Check what blocks node removal (PDBs, bare Pods, local storage) so that idle nodes do disappear.

  • Make cost visible per namespace or team using labels and a cost allocation tool, so that scaling decisions have owners.

11. Observability for Autoscaling

You cannot trust autoscaling that you cannot see. Build a dashboard and alerts around the whole chain. The metric names below assume kube-state-metrics, cAdvisor and the scheduler metrics are collected by Prometheus; verify the names in your environment.

What to watch

Example metric or command

Why

HPA desired vs current vs max replicas

kubehorizontalpodautoscalerstatusdesiredreplicas, statuscurrentreplicas, specmaxreplicas

Shows when the HPA is capped or lagging

Pending Pods

kubepodstatusphase{phase="Pending"}; scheduler pending Pods metric

Capacity or constraint problems

Node count and utilisation

kubenodeinfo, node CPU and memory metrics, requests vs allocatable

Shows headroom and node autoscaler activity

CPU throttling

containercpucfsthrottledperiodstotal over containercpucfsperiodstotal

Hidden latency from low CPU limits

OOM kills and restarts

kubepodcontainerstatuslastterminated_reason, restart counts

Sizing problems under load

Scale events

kubectl get events for SuccessfulRescale and TriggeredScaleUp

Audit trail of autoscaler decisions

Application SLIs

Latency, error rate, saturation, queue lag

The ultimate measure of whether scaling works


Useful alerts include: HPA at maxReplicas for more than a set time, Pending Pods for more than a few minutes, HPA unable to fetch metrics, node group at its maximum, sustained high latency with low replica counts, and a high rate of Pod restarts during scale events.

12. Testing Scaling: A Hands-On Drill

Test scaling before production traffic does it for you. This drill measures your own time budget on a staging cluster that resembles production. Use a non-production environment, because it generates load and may create nodes.

  1. Deploy the application with realistic requests, probes, an HPA and a PDB. Note the starting replica and node counts.

  2. Start a load test that ramps traffic in steps, using a load generator of your choice, while recording timestamps.

  3. Record the time at which: the metric crosses the target, the HPA raises the desired replicas, new Pods are created, Pods become Pending (if they do), nodes are added, Pods become Ready, and latency returns to normal.

  4. Stop the load and record when replicas and nodes scale in. Check for dropped requests during scale-in.

  5. Compare the measured time to new capacity with your traffic growth rate and adjust targets, headroom and minimums.

# Watch the chain during the test (each in its own terminal)

kubectl get hpa web -n shop -w

kubectl get pods -n shop -l app=web -w

kubectl get nodes -w

kubectl get events -n shop --sort-by=.lastTimestamp | tail -20

 

# Pending Pods and reasons

kubectl get pods -n shop --field-selector=status.phase=Pending

kubectl describe pod <pending-pod> -n shop | sed -n '/Events:/,$p'


Expected result: you obtain a timeline for your own environment, for example how many seconds from threshold to ready Pods with spare node capacity, and how many more when a node must be added. Repeat after each change, and rerun periodically, since images, dependencies and cluster settings change over time.

13. Failure Modes and Troubleshooting

Symptom

Likely cause

What to check or do

Latency rises during a spike, then recovers

Scale-out slower than traffic growth

Measure the path; add headroom, scheduled capacity or lower targets

HPA at maxReplicas, errors continue

Cap too low, or the bottleneck is elsewhere

Check dependency saturation before raising the cap

Pods Pending, nodes not added

Node group maximum, cloud quota, constraint no group satisfies

Check autoscaler events and logs, quota and node group limits

New Pods fail readiness

Slow start-up, missing dependencies, wrong probe

Check probe settings, logs, startup time; use startup probe

Errors when scaling in

Requests dropped during termination

Add preStop delay, handle SIGTERM, adjust termination grace period

Database errors after scale-out

Connection or load limit reached

Apply a connection budget, pooling and a lower maxReplicas

Flapping replicas

Spiky metric or short stabilisation

Tune behavior; smoother metric; longer scale-down window

Nodes never scale in

PDBs, bare Pods, local storage, annotations

Review scale-down blockers in the Cluster Autoscaler article

Costs rise without traffic growth

Inflated requests, high minimums, forgotten test workloads

Right-size, review minimums, check namespaces

Autoscaling behaves differently from the docs

Managed service defaults or limits

Check provider documentation and settings


14. Production Readiness Checklist

Area

Check

Signals

Scaling metric proven proportional to load by a load test; metrics pipeline highly available

Pod autoscaling

Requests set; HPA min and max justified; behavior tuned; no conflicting VPA or fixed replicas

Node autoscaling

Node groups designed; min and max set; quotas and IP space verified; scale-down blockers reviewed

Application

Stateless or safely shared state; correct readiness, startup and liveness probes; graceful shutdown tested

Availability

PDBs, topology spread across zones, sensible priorities

Dependencies

Connection budget calculated; pooling, caching and rate limiting in place; backpressure designed

Add-ons

DNS, ingress, mesh and logging scale with the workloads

Headroom and timing

Measured time-to-capacity compared with traffic growth; overprovisioning or scheduled capacity as needed

Cost

Requests right-sized; limits on maximums; cost visibility per team

Observability

Dashboards and alerts for the whole chain

Testing

Regular load tests and game days; results documented

Ownership

Each workload has a documented autoscaling owner and runbook


15. Best Practices

  • Treat scaling as an end-to-end property of the system, from signal to dependency, not as a feature of one controller.

  • Measure your time to new capacity and design around it, using headroom, scheduled minimums and lower targets.

  • Choose scaling signals that are proportional to demand and validate them with load tests.

  • Keep requests accurate; every autoscaler depends on them.

  • Make applications scale-friendly: stateless, quick to start, honest probes, graceful shutdown.

  • Derive maxReplicas from the capacity of dependencies, and design backpressure.

  • Scale the add-ons and verify cloud quotas before they become incidents.

  • Use PDBs, topology spread and priorities so that scaling and disruption do not reduce availability.

  • Keep cost visible and review minimums, maximums and idle capacity regularly.

  • Monitor the full chain, alert on caps and Pending Pods, and rehearse with load tests.

  • Avoid conflicting controllers on the same workload and signal.

16. Conclusion

Autoscaling in production is less about turning on the HPA, VPA and Cluster Autoscaler and more about understanding the path that capacity takes from a demand signal to a ready Pod, and about protecting everything that capacity touches. Measure how long that path takes, shorten it where you can, add headroom or scheduled capacity where you cannot, and make sure that applications, add-ons and dependencies are ready for the scale you allow. Observe the whole chain, test it regularly, and keep cost and ownership visible. When these pieces are in place, applications scale smoothly with demand, scale back down when it fades, and do so predictably, which is exactly what production needs.