Observability
Logs, metrics, and traces across service boundaries — correlation, SLOs, and debugging distributed failures without guessing.
You cannot operate what you cannot see. In microservices, local logs are not enough — you need a cross-service story of each request and workflow.
Three signals, different jobs
- Logs — what happened in detail (errors, decisions, payloads you are allowed to keep)
- Metrics — how the system behaves in aggregate (latency, error rate, saturation, lag)
- Traces — where time went across services for one request or saga step
Seniors instrument for questions they will ask at 3 a.m., not for vanity dashboards.
Correlate everything
Propagate a correlation / trace id across HTTP headers and message metadata. Without it, reconstructing a payment path across orchestration, fraud, and PSP adapters is archaeology.
Log that id on every meaningful line. Include business ids (payment id, merchant id) where safe — not secrets or full card data.
SLIs and SLOs beat vibes
Define service-level indicators that match user pain: successful checkout rate, capture latency, reconciliation freshness. Alert on burn against SLOs, not on every spike of CPU.
Consumer lag on Kafka is often a better signal than “pod restarted.”
Make failures diagnosable
When a dependency fails, emit:
- which dependency
- timeout vs error vs rejection
- whether you retried or degraded
Opaque 500 Internal Error with no context burns hours. Structured logs and span attributes pay for themselves on the first incident.
Privacy and volume
Observability can leak PII and drown you in noise. Sample high-cardinality traces, redact sensitive fields, and keep cardinality under control on metrics labels.
Senior checklist
- Can you follow one user action across all services in under five minutes?
- Do on-call engineers know which dashboard is authoritative?
- Are alerts actionable, or just noisy?
If debugging still starts with SSH and hope, the architecture is unfinished.