Debugging and troubleshooting
A senior path through Pending, CrashLoop, ImagePull, and networking faults — status before speculation.
Kubernetes failures are usually explained by status, events, and logs — in that order.
Pending Pods
Check: scheduling failures in Events, insufficient CPU/memory, affinity/taints, PVC unbound, missing RuntimeClass. kubectl describe pod beats guessing. Cluster autoscaler only helps if a node shape can satisfy the Pod.
CrashLoopBackOff / OOMKilled
Read previous container logs (--previous). Distinguish app panic, bad config, missing dependency, and memory limit too low. Raising limits without a root cause just delays the next OOM.
ImagePullBackOff
Registry auth, wrong digest/tag, rate limits, or network policy to the registry. Fix the pull story; do not “retry hope.”
Service has no Endpoints
Label selector mismatch, all Pods unready, or wrong namespace. Connect Service → Endpoints/EndpointSlice → Pod readiness before blaming DNS.
DNS and timeouts
CoreDNS latency or misconfigured ndots/search domains show up as flaky calls. Check from a debug Pod in the same namespace. East-west timeouts are often readiness or NetworkPolicy, not “Kubernetes is slow.”
Node NotReady
kubelet, runtime, disk pressure, network plugin. Drain only when safe; understand whether workloads have PDBs and replicas elsewhere.
Control plane symptoms
API slow, etcd latency, webhook timeouts. App restarts will not fix a sick API server.
Keep a short personal runbook: describe → events → logs → related objects (Deploy/Svc/Ingress/PVC). That order ends most pages faster than restarting nodes.