kubernetes

Kubernetes Vertical Pod Autoscaler Explained

By Shubhankar Tripathi • • 5 min read

Kubernetes Vertical Pod Autoscaler Explained

bitcodematrix.com | Kubernetes Series

1. Introduction

Almost every team struggles to choose the right CPU and memory requests. Set them too high and you pay for capacity nobody uses; set them too low and Pods are throttled, evicted or OOMKilled. Worse, the right numbers change as the application evolves. The Vertical Pod Autoscaler (VPA) addresses this by observing real usage and recommending, or automatically applying, better resource requests for your containers.

This article explains what the VPA is, how its three components work, how to write a VerticalPodAutoscaler object, what the update modes mean, how it differs from and coexists with the Horizontal Pod Autoscaler, and how to use it safely in production. It builds on the articles on requests and limits, QoS classes, resource management and the HPA.

After reading it you should be able to:

  • Explain what the VPA does and what it does not do.

  • Describe the Recommender, Updater and Admission Controller.

  • Choose between the Off, Initial, Recreate and Auto update modes.

  • Constrain recommendations with a resource policy.

  • Combine VPA with HPA safely and troubleshoot common problems.

2. What the VPA Is

The VPA is a Kubernetes add-on maintained in the kubernetes/autoscaler project. Unlike the HPA, it is not built into the core controller manager. You install its components in the cluster, and it adds a custom resource called VerticalPodAutoscaler (API group autoscaling.k8s.io). Some managed Kubernetes services offer it as an option or add-on, so check what your provider supports.

The VPA adjusts the resource requests of containers (and, optionally, the limits in proportion). It does not change the number of replicas. In the language of the previous article, the HPA scales out, and the VPA scales up.

Aspect

HPA

VPA

Changes

Number of replicas

CPU and memory requests (and optionally limits) per container

Built in

Yes

No, installed as an add-on

Applying changes

Adds or removes Pods without restarting existing ones

Usually evicts and recreates Pods (restart) to apply new values

Best fit

Stateless services with variable load

Right-sizing; workloads that cannot scale out, such as single-instance databases or batch jobs

Main risk

Flapping, slow scale-up

Restarts, recommendations that do not fit on any node


3. Architecture: Three Components

Recommender, updater and admission controller cooperating through the VPA object

Figure 1: How the VPA components cooperate

3.1 Recommender

The Recommender watches the resource usage of Pods through the metrics API and also loads some historical usage when it starts. It keeps statistics for each container and computes a recommendation using percentiles of usage plus a safety margin. For memory it also reacts to OOM events, increasing the recommendation after a container is OOM-killed. The result is written to the status of the VPA object as four values per container: lowerBound, target, upperBound and uncappedTarget.

Value

Meaning

target

The recommended request to apply

lowerBound

Minimum sensible request; going below it is likely to hurt

upperBound

Maximum sensible request; going above it wastes resources

uncappedTarget

The target before applying your minAllowed and maxAllowed policy


The Recommender needs time to gather data. Immediately after creating a VPA you may see no recommendation or one with low confidence, and recommendations improve over hours and days as it sees more usage patterns.

3.2 Updater

The Updater checks running Pods against the recommendation. If a Pod's current requests are outside the recommended range (below lowerBound or above upperBound), it evicts the Pod so that the controller recreates it. The Updater uses the eviction API, so it respects Pod Disruption Budgets. By default it will only evict Pods when enough replicas of the workload are running (a minimum of two replicas, which is configurable), to avoid taking down a single-replica service.

3.3 Admission Controller

The Admission Controller is a mutating webhook. When a Pod that matches a VPA is created, whether as a replacement for an evicted Pod or during any rollout, the webhook rewrites the container requests (and limits, if configured) to the recommended values before the Pod is stored. This is how new values reach the Pod. It is also why the VPA works at creation time even in modes where the Updater does not evict.

4. Update Modes

The updatePolicy.updateMode field decides whether and how the VPA applies recommendations.

Mode

Behaviour

Typical use

Off

Only computes recommendations. No Pods are changed.

Safest way to start; read recommendations and apply them yourself

Initial

Sets requests only when Pods are created. Running Pods are never evicted.

Avoids VPA-driven restarts; new values apply on normal rollouts

Recreate

Evicts Pods that are out of range and applies values on recreation.

Workloads that tolerate restarts

Auto

Currently behaves like Recreate. May use in-place updates when supported by the platform.

General automatic mode; default if no mode is set

InPlaceOrRecreate

Newer option that tries to resize Pods in place and falls back to eviction; availability depends on the VPA and Kubernetes versions

Where in-place resize is available; verify support first


Kubernetes itself has an in-place Pod resize feature, which lets container resources change without recreating the Pod. Its maturity depends on the version, and VPA integration with it is evolving. Check the documentation of both projects for your versions before depending on it.

5. Defining a VPA

5.1 Recommendation-only VPA

apiVersion: autoscaling.k8s.io/v1

kind: VerticalPodAutoscaler

metadata:

  name: web-vpa

  namespace: shop

spec:

  targetRef:

    apiVersion: apps/v1

    kind: Deployment

    name: web

  updatePolicy:

    updateMode: "Off"


Note that the value Off must be quoted in YAML so that it is read as a string, not as a boolean.

5.2 VPA with a resource policy

apiVersion: autoscaling.k8s.io/v1

kind: VerticalPodAutoscaler

metadata:

  name: web-vpa

  namespace: shop

spec:

  targetRef:

    apiVersion: apps/v1

    kind: Deployment

    name: web

  updatePolicy:

    updateMode: "Auto"

  resourcePolicy:

    containerPolicies:

      - containerName: app

        minAllowed:

          cpu: 100m

          memory: 128Mi

        maxAllowed:

          cpu: "2"

          memory: 2Gi

        controlledResources: ["cpu", "memory"]

        controlledValues: RequestsAndLimits

      - containerName: istio-proxy

        mode: "Off"


Field

Meaning

targetRef

The workload whose Pods are managed (Deployment, StatefulSet, DaemonSet, Job controller and so on)

minAllowed, maxAllowed

Bounds for recommendations; set maxAllowed below your largest node to avoid unschedulable Pods

controlledResources

Which resources the VPA may manage; for example only memory

controlledValues

RequestsAndLimits (default) scales limits in proportion to requests; RequestsOnly changes only requests and leaves limits as they are

mode: "Off" (per container)

Excludes one container, such as a sidecar, from VPA control


With RequestsAndLimits, the VPA keeps the ratio between the limit and the request that you originally defined. For example, if you set a request of 200Mi and a limit of 400Mi, a new request of 300Mi leads to a limit of 600Mi. Pods that have requests but no limits stay without limits. If you want limits to stay fixed, use RequestsOnly.

6. Reading the Recommendations

kubectl describe vpa web-vpa -n shop

 

# Status:

#   Recommendation:

#     Container Recommendations:

#       Container Name:  app

#       Lower Bound:

#         Cpu:     120m

#         Memory:  262144k

#       Target:

#         Cpu:     250m

#         Memory:  524288k

#       Uncapped Target:

#         Cpu:     250m

#         Memory:  524288k

#       Upper Bound:

#         Cpu:     1

#         Memory:  1Gi


The values above are illustrative. A practical way to use the Off mode is to review the target periodically, compare it with your current requests, and update your manifests accordingly through your normal change process. This gives data-driven right-sizing without automatic restarts. Several open-source tools build dashboards on top of VPA recommendations in this way.

# Compact view of all VPAs and their recommendations

kubectl get vpa -A

 

# Raw recommendation for scripting

kubectl get vpa web-vpa -n shop \

  -o jsonpath='{.status.recommendation.containerRecommendations[0].target}'


7. VPA and HPA Together

The HPA and the VPA both react to resource usage, which creates a conflict. If the HPA scales on CPU utilisation (a percentage of the request) and the VPA changes the CPU request, the two interfere: a larger request lowers measured utilisation, so the HPA scales in, then the VPA sees higher per-Pod usage and raises requests again, and so on. The Kubernetes documentation and the VPA project therefore advise against using both on the same CPU or memory metric for the same workload.

Combination

Guidance

HPA on CPU and VPA on CPU

Avoid; they conflict

HPA on a custom or external metric (for example requests per second) and VPA on CPU and memory

Works; the two use different signals

HPA on CPU and VPA in Off mode

Safe; use VPA only as a source of recommendations for your requests

VPA managing memory only and HPA on CPU

Possible with controlledResources, but test carefully


8. Interaction with Other Components

  • Scheduler and node size: a recommendation larger than any node cannot be scheduled. Always set maxAllowed to a value that fits your largest node.

  • Cluster Autoscaler: when the VPA raises requests and Pods no longer fit, the node autoscaler may add nodes, so VPA changes can affect cost.

  • Pod Disruption Budgets: the Updater uses eviction, so PDBs limit how many Pods it can restart at once.

  • ResourceQuota and LimitRange: recommendations must still satisfy namespace quotas and LimitRange bounds, or Pods will be rejected.

  • QoS class: changing requests and limits can change a Pod's QoS class when it is recreated (for example from Guaranteed to Burstable if requests and limits stop being equal). With RequestsAndLimits the ratio is kept, so equal values stay equal.

  • Runtimes with their own memory settings: the VPA sees container memory, not the language heap. A JVM with a fixed heap size will not shrink because the VPA wants a smaller request, so combine VPA with sensible runtime flags.

9. Hands-On Walkthrough

This exercise installs a recommendation-only VPA for a deliberately under-sized Deployment. It needs a test cluster with metrics-server and the VPA components installed (see the autoscaler project documentation for the supported installation method for your version).

  1. Deploy a Deployment with two replicas and very small CPU and memory requests.

  2. Create a VPA in Off mode targeting it.

  3. Generate some load and wait a few minutes.

  4. Read the recommendation, then switch to Auto and watch the Pods being recreated.

cat <<EOF | kubectl apply -f -

apiVersion: apps/v1

kind: Deployment

metadata:

  name: hamster

spec:

  replicas: 2

  selector:

    matchLabels:

      app: hamster

  template:

    metadata:

      labels:

        app: hamster

    spec:

      containers:

        - name: hamster

          image: registry.k8s.io/ubuntu-slim:0.14

          resources:

            requests:

              cpu: 50m

              memory: 50Mi

          command: ["/bin/sh"]

          args:

            - "-c"

            - "while true; do timeout 0.5s yes >/dev/null; sleep 0.5s; done"

---

apiVersion: autoscaling.k8s.io/v1

kind: VerticalPodAutoscaler

metadata:

  name: hamster-vpa

spec:

  targetRef:

    apiVersion: apps/v1

    kind: Deployment

    name: hamster

  updatePolicy:

    updateMode: "Off"

EOF

 

# After a few minutes

kubectl describe vpa hamster-vpa


The Deployment follows the style of the VPA project's own example, which burns CPU in short bursts. The image tag is an assumption, so substitute any small image that has a shell and the yes command. Expected result: the VPA shows a target CPU well above 50m. Now apply the recommendation automatically:

kubectl patch vpa hamster-vpa --type=merge \

  -p '{"spec":{"updatePolicy":{"updateMode":"Auto"}}}'

 

kubectl get pods -l app=hamster -w

kubectl get pod <new-pod> -o jsonpath='{.spec.containers[0].resources}'

 

# Clean up

kubectl delete vpa hamster-vpa

kubectl delete deployment hamster


Expected result (illustrative): after the Updater evicts the old Pods, the replacement Pods show the new, larger requests injected by the Admission Controller. Timing varies; the Updater acts at intervals and respects the minimum replica rule.

10. Troubleshooting

Symptom

Likely cause

What to check or do

VPA shows no recommendation

Recommender not running, no metrics, or too little history

Check the VPA component Pods and logs; verify kubectl top; wait several minutes

Recommendations exist but Pods are never updated

Mode is Off or Initial, fewer than the minimum replicas, PDB blocks eviction, or requests already within range

Check updateMode, replica count and PDB; compare requests with lowerBound and upperBound

Pods stay Pending after an update

Recommended request is larger than free capacity or any node

Set maxAllowed; check scheduler events and node autoscaler

Pod rejected at creation

Admission webhook failure, or quota and LimitRange violation

Check webhook availability and certificates; review quotas

Frequent restarts of Pods

Recreate or Auto with a workload whose usage swings widely

Use Initial or Off; widen bounds; apply recommendations through releases

Memory recommendation keeps growing

Memory leak in the application

Investigate the leak; do not just let the VPA raise limits

HPA and VPA fight

Both act on CPU or memory of the same workload

Separate their metrics, or run VPA in Off mode

Limits changed unexpectedly

controlledValues is RequestsAndLimits

Use RequestsOnly if limits must stay fixed


11. Best Practices

  • Start in Off mode and compare recommendations with your current values before enabling automation.

  • Always define minAllowed and maxAllowed so that recommendations stay within sensible and schedulable bounds.

  • Use at least two replicas, a Pod Disruption Budget and good readiness probes for workloads where the VPA may evict Pods.

  • Do not use VPA and HPA on the same CPU or memory metric for the same workload.

  • Exclude sidecars from VPA control, or give them their own policy.

  • Choose RequestsOnly if you want to keep fixed limits, and consider the QoS implications.

  • Treat VPA recommendations as evidence, not truth: validate with load tests, especially for spiky or batch workloads.

  • Monitor evictions, restarts, OOM kills and Pending Pods after enabling automatic updates.

  • Check supported versions and the maturity of in-place resizing before relying on it in production.

12. Conclusion

The Vertical Pod Autoscaler turns resource sizing from guesswork into a measured, continuously updated process. Its Recommender learns from real usage, its Updater and Admission Controller apply new values, and its policy fields keep the results within safe bounds. Most teams get the best return by starting in Off mode, using the recommendations to right-size their manifests, and only then enabling Initial or Auto for workloads that tolerate restarts. Pair it with the HPA by giving each one a different signal, and keep an eye on the interaction with quotas, node capacity and Pod Disruption Budgets. Used this way, the VPA reduces waste, prevents OOM kills and keeps your requests honest as your applications change.