Kubernetes Vertical Pod Autoscaler Explained
Kubernetes Vertical Pod Autoscaler Explained
bitcodematrix.com | Kubernetes Series
1. Introduction
Almost every team struggles to choose the right CPU and memory requests. Set them too high and you pay for capacity nobody uses; set them too low and Pods are throttled, evicted or OOMKilled. Worse, the right numbers change as the application evolves. The Vertical Pod Autoscaler (VPA) addresses this by observing real usage and recommending, or automatically applying, better resource requests for your containers.
This article explains what the VPA is, how its three components work, how to write a VerticalPodAutoscaler object, what the update modes mean, how it differs from and coexists with the Horizontal Pod Autoscaler, and how to use it safely in production. It builds on the articles on requests and limits, QoS classes, resource management and the HPA.
After reading it you should be able to:
Explain what the VPA does and what it does not do.
Describe the Recommender, Updater and Admission Controller.
Choose between the Off, Initial, Recreate and Auto update modes.
Constrain recommendations with a resource policy.
Combine VPA with HPA safely and troubleshoot common problems.
2. What the VPA Is
The VPA is a Kubernetes add-on maintained in the kubernetes/autoscaler project. Unlike the HPA, it is not built into the core controller manager. You install its components in the cluster, and it adds a custom resource called VerticalPodAutoscaler (API group autoscaling.k8s.io). Some managed Kubernetes services offer it as an option or add-on, so check what your provider supports.
The VPA adjusts the resource requests of containers (and, optionally, the limits in proportion). It does not change the number of replicas. In the language of the previous article, the HPA scales out, and the VPA scales up.
3. Architecture: Three Components
Figure 1: How the VPA components cooperate
3.1 Recommender
The Recommender watches the resource usage of Pods through the metrics API and also loads some historical usage when it starts. It keeps statistics for each container and computes a recommendation using percentiles of usage plus a safety margin. For memory it also reacts to OOM events, increasing the recommendation after a container is OOM-killed. The result is written to the status of the VPA object as four values per container: lowerBound, target, upperBound and uncappedTarget.
The Recommender needs time to gather data. Immediately after creating a VPA you may see no recommendation or one with low confidence, and recommendations improve over hours and days as it sees more usage patterns.
3.2 Updater
The Updater checks running Pods against the recommendation. If a Pod's current requests are outside the recommended range (below lowerBound or above upperBound), it evicts the Pod so that the controller recreates it. The Updater uses the eviction API, so it respects Pod Disruption Budgets. By default it will only evict Pods when enough replicas of the workload are running (a minimum of two replicas, which is configurable), to avoid taking down a single-replica service.
3.3 Admission Controller
The Admission Controller is a mutating webhook. When a Pod that matches a VPA is created, whether as a replacement for an evicted Pod or during any rollout, the webhook rewrites the container requests (and limits, if configured) to the recommended values before the Pod is stored. This is how new values reach the Pod. It is also why the VPA works at creation time even in modes where the Updater does not evict.
4. Update Modes
The updatePolicy.updateMode field decides whether and how the VPA applies recommendations.
Kubernetes itself has an in-place Pod resize feature, which lets container resources change without recreating the Pod. Its maturity depends on the version, and VPA integration with it is evolving. Check the documentation of both projects for your versions before depending on it.
5. Defining a VPA
5.1 Recommendation-only VPA
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: web-vpa
namespace: shop
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: web
updatePolicy:
updateMode: "Off"
Note that the value Off must be quoted in YAML so that it is read as a string, not as a boolean.
5.2 VPA with a resource policy
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: web-vpa
namespace: shop
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: web
updatePolicy:
updateMode: "Auto"
resourcePolicy:
containerPolicies:
- containerName: app
minAllowed:
cpu: 100m
memory: 128Mi
maxAllowed:
cpu: "2"
memory: 2Gi
controlledResources: ["cpu", "memory"]
controlledValues: RequestsAndLimits
- containerName: istio-proxy
mode: "Off"
With RequestsAndLimits, the VPA keeps the ratio between the limit and the request that you originally defined. For example, if you set a request of 200Mi and a limit of 400Mi, a new request of 300Mi leads to a limit of 600Mi. Pods that have requests but no limits stay without limits. If you want limits to stay fixed, use RequestsOnly.
6. Reading the Recommendations
kubectl describe vpa web-vpa -n shop
# Status:
# Recommendation:
# Container Recommendations:
# Container Name: app
# Lower Bound:
# Cpu: 120m
# Memory: 262144k
# Target:
# Cpu: 250m
# Memory: 524288k
# Uncapped Target:
# Cpu: 250m
# Memory: 524288k
# Upper Bound:
# Cpu: 1
# Memory: 1Gi
The values above are illustrative. A practical way to use the Off mode is to review the target periodically, compare it with your current requests, and update your manifests accordingly through your normal change process. This gives data-driven right-sizing without automatic restarts. Several open-source tools build dashboards on top of VPA recommendations in this way.
# Compact view of all VPAs and their recommendations
kubectl get vpa -A
# Raw recommendation for scripting
kubectl get vpa web-vpa -n shop \
-o jsonpath='{.status.recommendation.containerRecommendations[0].target}'
7. VPA and HPA Together
The HPA and the VPA both react to resource usage, which creates a conflict. If the HPA scales on CPU utilisation (a percentage of the request) and the VPA changes the CPU request, the two interfere: a larger request lowers measured utilisation, so the HPA scales in, then the VPA sees higher per-Pod usage and raises requests again, and so on. The Kubernetes documentation and the VPA project therefore advise against using both on the same CPU or memory metric for the same workload.
8. Interaction with Other Components
Scheduler and node size: a recommendation larger than any node cannot be scheduled. Always set maxAllowed to a value that fits your largest node.
Cluster Autoscaler: when the VPA raises requests and Pods no longer fit, the node autoscaler may add nodes, so VPA changes can affect cost.
Pod Disruption Budgets: the Updater uses eviction, so PDBs limit how many Pods it can restart at once.
ResourceQuota and LimitRange: recommendations must still satisfy namespace quotas and LimitRange bounds, or Pods will be rejected.
QoS class: changing requests and limits can change a Pod's QoS class when it is recreated (for example from Guaranteed to Burstable if requests and limits stop being equal). With RequestsAndLimits the ratio is kept, so equal values stay equal.
Runtimes with their own memory settings: the VPA sees container memory, not the language heap. A JVM with a fixed heap size will not shrink because the VPA wants a smaller request, so combine VPA with sensible runtime flags.
9. Hands-On Walkthrough
This exercise installs a recommendation-only VPA for a deliberately under-sized Deployment. It needs a test cluster with metrics-server and the VPA components installed (see the autoscaler project documentation for the supported installation method for your version).
Deploy a Deployment with two replicas and very small CPU and memory requests.
Create a VPA in Off mode targeting it.
Generate some load and wait a few minutes.
Read the recommendation, then switch to Auto and watch the Pods being recreated.
cat <<EOF | kubectl apply -f -
apiVersion: apps/v1
kind: Deployment
metadata:
name: hamster
spec:
replicas: 2
selector:
matchLabels:
app: hamster
template:
metadata:
labels:
app: hamster
spec:
containers:
- name: hamster
image: registry.k8s.io/ubuntu-slim:0.14
resources:
requests:
cpu: 50m
memory: 50Mi
command: ["/bin/sh"]
args:
- "-c"
- "while true; do timeout 0.5s yes >/dev/null; sleep 0.5s; done"
---
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: hamster-vpa
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: hamster
updatePolicy:
updateMode: "Off"
EOF
# After a few minutes
kubectl describe vpa hamster-vpa
The Deployment follows the style of the VPA project's own example, which burns CPU in short bursts. The image tag is an assumption, so substitute any small image that has a shell and the yes command. Expected result: the VPA shows a target CPU well above 50m. Now apply the recommendation automatically:
kubectl patch vpa hamster-vpa --type=merge \
-p '{"spec":{"updatePolicy":{"updateMode":"Auto"}}}'
kubectl get pods -l app=hamster -w
kubectl get pod <new-pod> -o jsonpath='{.spec.containers[0].resources}'
# Clean up
kubectl delete vpa hamster-vpa
kubectl delete deployment hamster
Expected result (illustrative): after the Updater evicts the old Pods, the replacement Pods show the new, larger requests injected by the Admission Controller. Timing varies; the Updater acts at intervals and respects the minimum replica rule.
10. Troubleshooting
11. Best Practices
Start in Off mode and compare recommendations with your current values before enabling automation.
Always define minAllowed and maxAllowed so that recommendations stay within sensible and schedulable bounds.
Use at least two replicas, a Pod Disruption Budget and good readiness probes for workloads where the VPA may evict Pods.
Do not use VPA and HPA on the same CPU or memory metric for the same workload.
Exclude sidecars from VPA control, or give them their own policy.
Choose RequestsOnly if you want to keep fixed limits, and consider the QoS implications.
Treat VPA recommendations as evidence, not truth: validate with load tests, especially for spiky or batch workloads.
Monitor evictions, restarts, OOM kills and Pending Pods after enabling automatic updates.
Check supported versions and the maturity of in-place resizing before relying on it in production.
12. Conclusion
The Vertical Pod Autoscaler turns resource sizing from guesswork into a measured, continuously updated process. Its Recommender learns from real usage, its Updater and Admission Controller apply new values, and its policy fields keep the results within safe bounds. Most teams get the best return by starting in Off mode, using the recommendations to right-size their manifests, and only then enabling Initial or Auto for workloads that tolerate restarts. Pair it with the HPA by giving each one a different signal, and keep an eye on the interaction with quotas, node capacity and Pod Disruption Budgets. Used this way, the VPA reduces waste, prevents OOM kills and keeps your requests honest as your applications change.