Learning · Kubernetes

Day-2 operations

GitOps, upgrades, backups, and cost — running the platform after the first demo cluster.

The hard part of Kubernetes starts after the first green Deployment.

Declare desired state

Prefer GitOps or equivalent pipelines: cluster state reviewed in PRs, applied by controllers (Argo CD, Flux, or hardened CI). Snowflake kubectl changes drift and cannot be audited.

Upgrade deliberately

Control plane, nodes, and add-ons (CNI, Ingress, metrics) have compatibility matrices. Drain with PDBs, surge capacity, and a rollback story. Skipping minor versions casually is how you earn a weekend.

Backup what you cannot rebuild

etcd (or managed control-plane backups), persistent volumes you care about, and cluster config in git. Practice restore. Stateless apps reconstruct from images; data does not.

Cost and idle waste

Over-requests, idle LoadBalancers, orphaned PVCs, and always-on non-prod replicas dominate bills. Right-size from metrics; turn down environments that sleep; delete unused Ingresses.

Platform vs product ownership

Platform owns cluster health, defaults, and paved roads. Product owns service charts, SLOs, and on-call for their apps. Blurring that line produces tickets that bounce forever.

Day-2 success looks boring: predictable deploys, tested upgrades, and clear ownership when something pages.

← Kubernetes