Autoscaling
HPA, VPA, and cluster autoscaler — scaling Pods and nodes without thrash or pending queues.
Autoscaling is three different problems: scale replicas, resize Pod resources, and grow nodes.
Horizontal Pod Autoscaler
Scales Deployment/StatefulSet replicas from CPU, memory, or custom metrics (queue depth, QPS, lag). Needs:
- honest resource requests
- metrics that track user pain
- stabilization so you do not flap
HPA cannot fix a single-threaded hotspot or a sync call chain. Scale the bottleneck you measured.
Vertical Pod Autoscaler
Recommends or applies request/limit changes. Useful for rightsizing; dangerous if it restarts Pods mid-traffic without a plan. Prefer recommendation mode first; apply carefully for Guaranteed QoS services.
Cluster autoscaler / node pools
Adds/removes nodes when Pods are unschedulable or nodes underutilized. Respects taints, labels, and PDB constraints. Pending Pods with impossible affinity will not magically get nodes that fit.
Interaction effects
HPA up → more Pods Pending → cluster autoscaler adds nodes → bill rises. Without max replicas and budget alerts, scale-out is an open wallet. Set ceilings.
Cron and event-driven scale
Scheduled scaling (business hours) and KEDA-style event scalers help bursty workloads. Still define min replicas for sudden traffic after idle.
Smell test
If autoscaling is “on” but on-call still manually kubectl scale, the signals or limits are wrong — fix metrics and policies, do not add another dashboard.