Learning · Kubernetes

Resources, QoS, and quotas

Requests, limits, QoS classes, and namespace quotas — fair sharing without silent eviction surprises.

CPU and memory settings are scheduling and reliability contracts. Get them wrong and you get Pending, throttling, or OOMKills.

Requests vs limits

  • Request — what the scheduler reserves; baseline for bin-packing
  • Limit — hard cap (memory) or throttle ceiling (CPU)

Memory limit breached → OOMKill. CPU limit breached → throttle (latency spikes). Many teams set memory request=limit for Guaranteed QoS on latency-critical services.

QoS classes

  • Guaranteed — request=limit for all containers
  • Burstable — requests set, limits higher or mixed
  • BestEffort — no requests/limits; first to evict under node pressure

Know which class your Pod is; eviction order follows it under memory pressure.

LimitRanges and ResourceQuotas

Namespace defaults and caps stop unbounded pods and runaway teams. Platform sets sane defaults; products tune from metrics (VPA suggestions as input, not blind autopilot).

Ephemeral storage

Root filesystem and emptyDir can fill and evict Pods. Log to sidecars/shippers with rotation; do not treat the container root as infinite disk.

Rightsizing loop

Start from load tests and production metrics. Revisit after feature launches. Fantasy requests are how clusters feel “full” while CPU idle looks high in aggregate.

← Kubernetes