Kubernetes Autoscaling: How Applications Scale in Production
Kubernetes Autoscaling: How Applications Scale in Production
bitcodematrix.com | Kubernetes Series
1. Introduction
Enabling an autoscaler is the easy part of scaling in production. The hard part is making sure that when load arrives, new capacity is ready quickly enough, the application can use it safely, the databases and downstream services behind it can absorb the extra traffic, and the bill stays under control. Many "autoscaling" incidents are not caused by the autoscalers at all: the new Pods take four minutes to become ready, the database runs out of connections, the nodes cannot be provisioned because a cloud quota was reached, or the metric that drives scaling does not reflect real demand.
This article is the capstone of the autoscaling part of this series. The earlier articles explain the HPA, the VPA and the Cluster Autoscaler individually and compare them. Here we look at the whole system from the production point of view: the scale-out path and its time budget, choosing scaling signals, preparing applications and platform, protecting dependencies, controlling cost, observing and testing scaling, and handling failure modes.
After reading it you should be able to:
Describe every step between a traffic spike and new Pods serving requests, and estimate how long it takes.
Choose scaling signals and strategies (reactive, scheduled, event-driven) for different workloads.
Prepare applications (probes, graceful shutdown, pools) and the platform (headroom, quotas, images) for scaling.
Protect databases and downstream services from the load of a scale-out.
Monitor, load test and troubleshoot autoscaling, and apply a production readiness checklist.
2. Scaling Is a Chain, Not a Switch
For a request to be served by a new replica, a whole chain of things must work. Autoscalers handle only part of it.
The weakest link sets the real scaling speed and safety. The rest of this article works through the chain.
3. Scaling Strategies
There are three broad strategies, and mature platforms mix them.
Reactive scaling alone always lags demand by the time it takes to measure, decide, schedule and start. When a peak is known in advance, combine reactive scaling with a higher minimum during the peak window so that the reactive part only handles the surprise on top. An event-driven autoscaler such as KEDA can express this with a cron trigger. The following sketch is illustrative; check the documentation of the version you use.
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
name: web-business-hours
namespace: shop
spec:
scaleTargetRef:
name: web
minReplicaCount: 3
maxReplicaCount: 30
triggers:
- type: cron
metadata:
timezone: Asia/Kolkata
start: "0 8 1-5"
end: "0 20 1-5"
desiredReplicas: "10"
- type: cpu
metricType: Utilization
metadata:
value: "65"
Here the workload keeps at least ten replicas during weekday business hours, and the CPU trigger can add more on top. Outside those hours it falls back to the minimum. Do not combine a KEDA ScaledObject with a separate HPA on the same workload; KEDA manages its own HPA.
4. The Scale-Out Path and Its Time Budget
Figure 1: The scale-out path and where the time goes
The numbers are examples, not promises. Measure your own path with a test (section 12). The practical conclusion is that you need to compare your measured time to new capacity with how fast your traffic can grow. If a spike can double traffic in 30 seconds and new capacity takes five minutes, you need headroom, scheduled capacity, or a lower target utilisation, not a faster HPA.
4.1 Ways to shorten the path
Lower the HPA target (for example 50 to 65 percent instead of 80) so that spare capacity exists while new Pods start.
Keep node headroom with low-priority overprovisioning Pods so that new Pods can be scheduled immediately.
Use small images and keep them in a registry close to the cluster; consider pre-pulling or image caching features of your platform.
Make the application start quickly: avoid heavy initialisation, defer warm-up, and use startup probes for slow starters.
Tune HPA behavior to allow fast scale-up (large step sizes, no stabilisation delay) and slower scale-down.
Right-size requests so that fewer Pods are needed and each fits easily.
5. Choosing Scaling Signals
The HPA is only as good as the metric it watches. A good signal rises when demand rises, falls when demand falls, is available quickly, and is proportional (doubling the load roughly doubles the metric per Pod).
When possible, derive the target from a load test: find the load per Pod at which latency starts to degrade, then set the target to roughly 60 to 70 percent of that. Treat these percentages as a starting point rather than a rule.
6. Preparing the Application
6.1 Stateless and horizontally safe
Horizontal scaling works only if replicas are interchangeable. Keep session state in a shared store or in signed tokens, avoid local files that other replicas need, and make background jobs safe to run on several replicas (leader election or locking where needed).
6.2 Probes that tell the truth
Readiness probe: returns success only when the Pod can serve real traffic. Without it, new Pods receive traffic too early and users see errors during scale-out.
Startup probe: protects slow-starting applications from being killed by liveness checks while they start.
Liveness probe: use carefully; an overly strict liveness probe under high load can cause restarts that make the load problem worse.
6.3 Graceful shutdown for scale-in
Scale-in and node drains terminate Pods. To avoid dropping requests, the Pod should stop receiving new traffic, finish current work, and exit. Kubernetes sends SIGTERM, and removal from Service endpoints happens in parallel, so there can be a short period in which traffic still arrives. A short preStop delay and a proper SIGTERM handler address this.
spec:
terminationGracePeriodSeconds: 60
containers:
- name: app
image: registry.example.com/shop/web:1.0
lifecycle:
preStop:
exec:
command: ["sleep", "10"]
readinessProbe:
httpGet:
path: /healthz
port: 8080
periodSeconds: 5
startupProbe:
httpGet:
path: /healthz
port: 8080
failureThreshold: 30
periodSeconds: 2
The preStop sleep gives load balancers and endpoint controllers time to stop sending traffic before the application starts shutting down, and the application itself should drain connections when it receives SIGTERM. The sleep command requires the container image to include it. The image name and values are placeholders to adapt.
6.4 Requests, spreading and disruption budgets
Set realistic CPU and memory requests (see the article on requests and limits).
Use topology spread constraints so that replicas are spread across nodes and zones.
Use a Pod Disruption Budget so that voluntary disruptions, including node removal by the autoscaler, never take down too many replicas at once.
Choose a QoS class that matches the workload's importance.
7. Preparing the Platform
A frequent hidden bottleneck is the add-ons. If the application scales to 200 replicas but CoreDNS, the ingress controller or the service mesh gateway stays at its original size, the add-on becomes the new limit.
8. Protecting Dependencies
Scaling out the application multiplies its demands on everything behind it. This is the most common production surprise: the web tier scales perfectly and takes the database down.
A simple connection budget calculation illustrates the point:
maxReplicas x connections per Pod <= database max connections - reserved connections
Example: 30 replicas x 10 connections = 300 connections
database allows 400, reserve 100 for admin and jobs
300 <= 300, so maxReplicas of 30 is the upper bound
Treat the HPA maxReplicas as a safety control derived from the capacity of the whole system, not just a cost cap. Add backpressure (queues, load shedding, timeouts) so that overload degrades gracefully instead of cascading.
9. Patterns by Workload Type
10. Cost Control
Autoscaling reduces cost by removing idle capacity, but it can also hide waste or add cost if configured carelessly.
Right-size requests. Inflated requests make the HPA add Pods early and make nodes look full. Use VPA recommendations or usage data.
Set maxReplicas and node group maximums as budget and safety limits, and alert when they are reached.
Choose minimums deliberately. A high minimum is a fixed cost; a low one is a risk. Review them against real traffic.
Use cheaper capacity where it is safe, such as spot or preemptible nodes for fault-tolerant workloads, with PDBs and multiple instance types.
Allow scale-down. Check what blocks node removal (PDBs, bare Pods, local storage) so that idle nodes do disappear.
Make cost visible per namespace or team using labels and a cost allocation tool, so that scaling decisions have owners.
11. Observability for Autoscaling
You cannot trust autoscaling that you cannot see. Build a dashboard and alerts around the whole chain. The metric names below assume kube-state-metrics, cAdvisor and the scheduler metrics are collected by Prometheus; verify the names in your environment.
Useful alerts include: HPA at maxReplicas for more than a set time, Pending Pods for more than a few minutes, HPA unable to fetch metrics, node group at its maximum, sustained high latency with low replica counts, and a high rate of Pod restarts during scale events.
12. Testing Scaling: A Hands-On Drill
Test scaling before production traffic does it for you. This drill measures your own time budget on a staging cluster that resembles production. Use a non-production environment, because it generates load and may create nodes.
Deploy the application with realistic requests, probes, an HPA and a PDB. Note the starting replica and node counts.
Start a load test that ramps traffic in steps, using a load generator of your choice, while recording timestamps.
Record the time at which: the metric crosses the target, the HPA raises the desired replicas, new Pods are created, Pods become Pending (if they do), nodes are added, Pods become Ready, and latency returns to normal.
Stop the load and record when replicas and nodes scale in. Check for dropped requests during scale-in.
Compare the measured time to new capacity with your traffic growth rate and adjust targets, headroom and minimums.
# Watch the chain during the test (each in its own terminal)
kubectl get hpa web -n shop -w
kubectl get pods -n shop -l app=web -w
kubectl get nodes -w
kubectl get events -n shop --sort-by=.lastTimestamp | tail -20
# Pending Pods and reasons
kubectl get pods -n shop --field-selector=status.phase=Pending
kubectl describe pod <pending-pod> -n shop | sed -n '/Events:/,$p'
Expected result: you obtain a timeline for your own environment, for example how many seconds from threshold to ready Pods with spare node capacity, and how many more when a node must be added. Repeat after each change, and rerun periodically, since images, dependencies and cluster settings change over time.
13. Failure Modes and Troubleshooting
14. Production Readiness Checklist
15. Best Practices
Treat scaling as an end-to-end property of the system, from signal to dependency, not as a feature of one controller.
Measure your time to new capacity and design around it, using headroom, scheduled minimums and lower targets.
Choose scaling signals that are proportional to demand and validate them with load tests.
Keep requests accurate; every autoscaler depends on them.
Make applications scale-friendly: stateless, quick to start, honest probes, graceful shutdown.
Derive maxReplicas from the capacity of dependencies, and design backpressure.
Scale the add-ons and verify cloud quotas before they become incidents.
Use PDBs, topology spread and priorities so that scaling and disruption do not reduce availability.
Keep cost visible and review minimums, maximums and idle capacity regularly.
Monitor the full chain, alert on caps and Pending Pods, and rehearse with load tests.
Avoid conflicting controllers on the same workload and signal.
16. Conclusion
Autoscaling in production is less about turning on the HPA, VPA and Cluster Autoscaler and more about understanding the path that capacity takes from a demand signal to a ready Pod, and about protecting everything that capacity touches. Measure how long that path takes, shorten it where you can, add headroom or scheduled capacity where you cannot, and make sure that applications, add-ons and dependencies are ready for the scale you allow. Observe the whole chain, test it regularly, and keep cost and ownership visible. When these pieces are in place, applications scale smoothly with demand, scale back down when it fades, and do so predictably, which is exactly what production needs.