Networking and exposure
Services, Ingress, and DNS — getting traffic to pods without turning the mesh into the product.
Pods have ephemeral IPs. Stable reachability comes from Services and whatever fronts them.
Service types with intent
- ClusterIP — internal east-west traffic (default mental model for microservices)
- NodePort / LoadBalancer — external entry when the platform expects it
- headless — when clients need pod DNS directly (StatefulSets, some meshes)
Pick one exposure path per use case. Publishing every Deployment as a public LoadBalancer is how cloud bills and attack surface grow.
Ingress / gateway is a product edge
TLS termination, routing, and auth at the edge belong in a deliberate Ingress/Gateway design — not ad-hoc per team with conflicting annotations. Centralize what must be consistent (TLS, WAF, request IDs); allow teams autonomy behind ClusterIP.
DNS and retries
kube-dns/CoreDNS failures look like “the app is down.” Clients need timeouts, bounded retries, and readiness-aware endpoints. Retry storms across a dependency mesh amplify outages — the same rules as microservice communication apply.
NetworkPolicy is allowlisting
Default-open pod networks are convenient and dangerous. Start with deny-ish policies in sensitive namespaces and open what the service needs. Policies only help if the CNI enforces them — verify on your platform.
Service mesh is optional complexity
mTLS, retries, and traffic splits can live in a mesh — or in app libraries and gateways. Add a mesh when the operational model is clear; do not install one to postpone fixing timeouts and authz in the services themselves.