Learning · Kubernetes

Autoscaling

HPA, VPA, and cluster autoscaler — scaling Pods and nodes without thrash or pending queues.

Autoscaling is three different problems: scale replicas, resize Pod resources, and grow nodes.

Horizontal Pod Autoscaler

Scales Deployment/StatefulSet replicas from CPU, memory, or custom metrics (queue depth, QPS, lag). Needs:

  • honest resource requests
  • metrics that track user pain
  • stabilization so you do not flap

HPA cannot fix a single-threaded hotspot or a sync call chain. Scale the bottleneck you measured.

Vertical Pod Autoscaler

Recommends or applies request/limit changes. Useful for rightsizing; dangerous if it restarts Pods mid-traffic without a plan. Prefer recommendation mode first; apply carefully for Guaranteed QoS services.

Cluster autoscaler / node pools

Adds/removes nodes when Pods are unschedulable or nodes underutilized. Respects taints, labels, and PDB constraints. Pending Pods with impossible affinity will not magically get nodes that fit.

Interaction effects

HPA up → more Pods Pending → cluster autoscaler adds nodes → bill rises. Without max replicas and budget alerts, scale-out is an open wallet. Set ceilings.

Cron and event-driven scale

Scheduled scaling (business hours) and KEDA-style event scalers help bursty workloads. Still define min replicas for sudden traffic after idle.

Smell test

If autoscaling is “on” but on-call still manually kubectl scale, the signals or limits are wrong — fix metrics and policies, do not add another dashboard.

← Kubernetes