Learning · Kubernetes

Debugging and troubleshooting

A senior path through Pending, CrashLoop, ImagePull, and networking faults — status before speculation.

Kubernetes failures are usually explained by status, events, and logs — in that order.

Pending Pods

Check: scheduling failures in Events, insufficient CPU/memory, affinity/taints, PVC unbound, missing RuntimeClass. kubectl describe pod beats guessing. Cluster autoscaler only helps if a node shape can satisfy the Pod.

CrashLoopBackOff / OOMKilled

Read previous container logs (--previous). Distinguish app panic, bad config, missing dependency, and memory limit too low. Raising limits without a root cause just delays the next OOM.

ImagePullBackOff

Registry auth, wrong digest/tag, rate limits, or network policy to the registry. Fix the pull story; do not “retry hope.”

Service has no Endpoints

Label selector mismatch, all Pods unready, or wrong namespace. Connect Service → Endpoints/EndpointSlice → Pod readiness before blaming DNS.

DNS and timeouts

CoreDNS latency or misconfigured ndots/search domains show up as flaky calls. Check from a debug Pod in the same namespace. East-west timeouts are often readiness or NetworkPolicy, not “Kubernetes is slow.”

Node NotReady

kubelet, runtime, disk pressure, network plugin. Drain only when safe; understand whether workloads have PDBs and replicas elsewhere.

Control plane symptoms

API slow, etcd latency, webhook timeouts. App restarts will not fix a sick API server.

Keep a short personal runbook: describe → events → logs → related objects (Deploy/Svc/Ingress/PVC). That order ends most pages faster than restarting nodes.

← Kubernetes