Day-2 operations
GitOps, upgrades, backups, and cost — running the platform after the first demo cluster.
The hard part of Kubernetes starts after the first green Deployment.
Declare desired state
Prefer GitOps or equivalent pipelines: cluster state reviewed in PRs, applied by controllers (Argo CD, Flux, or hardened CI). Snowflake kubectl changes drift and cannot be audited.
Upgrade deliberately
Control plane, nodes, and add-ons (CNI, Ingress, metrics) have compatibility matrices. Drain with PDBs, surge capacity, and a rollback story. Skipping minor versions casually is how you earn a weekend.
Backup what you cannot rebuild
etcd (or managed control-plane backups), persistent volumes you care about, and cluster config in git. Practice restore. Stateless apps reconstruct from images; data does not.
Cost and idle waste
Over-requests, idle LoadBalancers, orphaned PVCs, and always-on non-prod replicas dominate bills. Right-size from metrics; turn down environments that sleep; delete unused Ingresses.
Platform vs product ownership
Platform owns cluster health, defaults, and paved roads. Product owns service charts, SLOs, and on-call for their apps. Blurring that line produces tickets that bounce forever.
Day-2 success looks boring: predictable deploys, tested upgrades, and clear ownership when something pages.