Kubernetes Autoscaling
Autoscaling on Kubernetes happens at two layers. Pod autoscaling changes how many replicas a workload runs, or how big each one is. Node autoscaling changes how many machines the cluster has so those Pods have somewhere to run. The two work together: the Horizontal Pod Autoscaler adds Pods, some of them can't be scheduled, and the node autoscaler adds a node for them.
Done well, autoscaling keeps latency flat during spikes and cuts cloud costs during quiet hours. Done badly, it flaps, lags behind traffic, or scales on a metric that doesn't reflect load at all.
TL;DR
- HPA (Horizontal Pod Autoscaler) adds or removes replicas based on CPU, memory, custom, or external metrics. It's the default choice for stateless services.
- VPA (Vertical Pod Autoscaler) recommends or sets CPU/memory requests. Use it for right-sizing, and don't combine it with an HPA on the same resource metric.
- KEDA scales on event sources (queue depth, Kafka lag, cron, Prometheus queries) and can scale to zero.
- Cluster Autoscaler adds or removes nodes from predefined node groups when Pods are pending or nodes are underused.
- Karpenter provisions right-sized nodes directly from pending Pod requirements, which is faster and more flexible on AWS and increasingly elsewhere.
- Autoscaling only works if resource requests are accurate; HPA utilization is measured against requests.
Quick Example
Scale an API between 3 and 30 replicas to hold average CPU near 70% of requests, with calmer scale-down:
Scale-up can double the replica count every 30 seconds. Scale-down waits five minutes of sustained low load, then removes at most two Pods a minute.
Core Concepts
Horizontal Pod Autoscaler
The HPA controller runs a loop (every 15 seconds by default) that computes:
Metric types:
Custom and external metrics need an adapter such as the Prometheus Adapter or a cloud metrics adapter, or KEDA, which serves the external metrics API for you. With several metrics configured, the HPA scales to whichever yields the most replicas.
Vertical Pod Autoscaler
The VPA watches actual usage and recommends requests. In Off mode it only publishes recommendations (the safe way to start). In Initial mode it sets requests at Pod creation. In Recreate/Auto mode it evicts Pods to apply new sizes; with in-place Pod resize (now GA), some changes can apply without a restart. VPA is great for right-sizing batch jobs and singletons, and for finding sensible requests for everything else.
KEDA
KEDA (Kubernetes Event-Driven Autoscaling) adds a ScaledObject resource with dozens of scalers: Kafka consumer lag, RabbitMQ, SQS, Azure Service Bus, Redis lists, Prometheus queries, cron windows, and more. It drives an HPA under the hood and adds scale-to-zero: with no events, the workload runs zero replicas, and KEDA scales it back up when messages arrive. It's the natural fit for message queue consumers and background jobs.
Node Autoscaling
- Cluster Autoscaler works with node groups (ASGs, managed instance groups, VM scale sets). When Pods are pending for lack of capacity, it grows the group that would fit them. When a node's Pods could fit elsewhere, it drains and removes the node.
- Karpenter skips node groups. It looks at pending Pods' requirements (CPU, memory, architecture, zone, spot vs on-demand) and launches exactly the instance types that fit, then consolidates by replacing underused nodes with cheaper ones. It's generally faster and cheaper, at the cost of a different operating model.
Choosing What to Scale On
- CPU works for compute-bound, request-driven services. It's simple and usually good enough.
- Memory is rarely a good HPA signal, since most runtimes don't release memory when load drops.
- Requests per second or concurrency track load directly for web services. Scale on in-flight requests per Pod when latency matters most.
- Queue depth or lag is the right signal for consumers. CPU on an idle-but-backlogged worker says nothing.
- Latency is tempting but a lagging, noisy signal. Use it for alerting, not as the primary scaling metric.
Best Practices
Get Requests Right First
HPA utilization is a percentage of requests. Requests set far too high means the HPA never scales; far too low means it scales constantly. Run VPA in recommendation mode or review real usage before tuning HPA targets.
Leave Headroom and Tune Behavior
Target 60–75% CPU, not 95%. New Pods take time to start and warm up, and node provisioning can take a minute or more. Scale up aggressively and scale down slowly with a stabilization window to avoid flapping.
Keep a Sensible Minimum
minReplicas should cover baseline traffic and survive losing a zone. Scale-to-zero suits async workers, not latency-sensitive APIs, because the first request pays for a cold start.
Budget Disruptions
Node consolidation drains nodes. PodDisruptionBudgets keep enough replicas running while it happens, and karpenter.sh/do-not-disrupt or safe-to-evict annotations protect long-running jobs.
Common Mistakes
HPA and VPA Fighting Over CPU
If you use both, have the HPA scale on a custom metric such as RPS, and let the VPA manage only resources the HPA doesn't use.
No Resource Requests at All
Without CPU requests, the HPA can't compute utilization and reports <unknown>. Nothing scales.
Scaling Pods Without Scaling Nodes
An HPA that creates Pods the cluster can't schedule just produces a pile of Pending Pods. Pair Pod autoscaling with Cluster Autoscaler or Karpenter, and set maxReplicas within what your node limits and budget allow.
FAQ
How fast does the HPA react?
The controller evaluates every 15 seconds, metrics-server scrapes about every 15 seconds, and new Pods then need to start and pass readiness. Expect 30–90 seconds for Pods to add capacity, plus node provisioning time if new nodes are needed. For predictable spikes, pre-scale on a schedule (KEDA's cron scaler works well).
Should I use Cluster Autoscaler or Karpenter?
On AWS, Karpenter is the common choice for new clusters: it provisions faster, picks cheaper instance types, and consolidates automatically. Cluster Autoscaler is mature, works on every major cloud, and fits well if you already manage node groups carefully. Managed options like GKE Autopilot and EKS Auto Mode handle node scaling for you.
Can Kubernetes scale to zero?
Not with the built-in HPA, whose minimum is one replica. KEDA and Knative both scale to zero and back up on demand. It's ideal for queue workers and low-traffic internal services.
Why isn't my HPA scaling?
Run kubectl describe hpa. Common causes: metrics-server isn't installed, Pods lack resource requests, the custom metrics adapter isn't serving the metric, or the workload is already at maxReplicas.
Related Topics
- Kubernetes — The platform overview
- Kubernetes Workloads — Requests, limits, and PodDisruptionBudgets
- Scalability — Horizontal vs vertical scaling in general
- Capacity Planning — Sizing for peak load
- Cloud Costs — Autoscaling as a cost lever
- Message Queues — Scaling consumers on backlog