Learning · Software Engineering Manifesto
Observability is part of done
Structured logs, metrics, and traces ship with the feature — not as a follow-up ticket after the outage.
Principle
If you cannot see how a change behaves in production, you did not finish it. Observability is a definition-of-done item for service work.
Therefore we practice
- Emit structured logs (fields, not only prose) with correlation/causation ids across hops
- Expose golden signals for the journey: rate, errors, latency (and saturation where relevant)
- Propagate trace context through sync and async boundaries
- Log outcomes at decision points (authorize, capture, route) — not only stack traces
- Dashboards and alerts exist before or with the rollout, not “later”
Smells
log.info("error " + e)with no ids, no structure, no context- Only CPU/memory charts when users report failed checkouts
- New endpoints with zero metrics
- Debugging by SSH into a random replica