Reliability and scaling
HPA, PDBs, disruption, and graceful shutdown — staying up when nodes and deploys move under you.
Kubernetes moves pods. Reliable services expect that.
Horizontal scaling
HPA scales replicas from CPU, memory, or custom metrics. It needs honest resource requests and metrics that track user pain (queue depth, lag, p95 latency) — not vanity CPU that never correlates.
Autoscaling cannot fix synchronous call chains or a single hot shard in a dependency. Scale the bottleneck you measured.
PodDisruptionBudgets
PDBs limit concurrent voluntary disruptions (drains, upgrades). Without them, a cluster upgrade can evict an entire Deployment at once. Set budgets that match your true redundancy — minAvailable: 1 on a singleton is a fiction.
Graceful shutdown
On termination: stop readiness, drain in-flight work, then exit. Align terminationGracePeriodSeconds with real drain time. Ignoring SIGTERM causes dropped requests during every deploy.
Quotas and limits at namespace level
ResourceQuotas and LimitRanges stop one team from scheduling the cluster into failure. Platform teams should set defaults; product teams should know their caps.
Multi-AZ thinking
Spread replicas across failure domains (topologySpreadConstraints / anti-affinity) when availability matters. Three replicas on one node is one failure away from zero.