Observability on Kubernetes
Logs, metrics, and traces in a world where pods disappear — correlating deploys with user pain.
Pods are temporary. Telemetry must outlive them.
Logs leave the node
Ship container logs to a central store (ELK, Loki, cloud logging). Include pod, namespace, deployment, and version labels. kubectl logs is for debugging a moment — not for retention or multi-replica search.
Metrics that map to SLOs
Watch request rate, errors, and latency at the Service/Ingress layer and inside the app. Add Kubernetes signals: restart counts, OOMKills, pending pods, PVC pressure, HPA replica count, and node conditions.
Alert on user impact first; alert on “pod restarted once” second.
Traces across hops
Propagate trace/correlation ids through Ingress and east-west calls. Without them, a timeout in service B during a rolling deploy is folklore.
Events and kubectl describe
Kubernetes Events explain scheduling failures, probe kills, and image pull errors. They are short-lived — scrape or archive important ones if you need post-incident evidence.
Link deploys to telemetry
Annotate releases (image digest, change ticket). When p99 spikes, you want “what shipped?” in the same timeline as “what broke?”
Observability on Kubernetes is the same discipline as microservices — with stronger labeling, because the instance you SSH’d to yesterday is gone.