Learning · Software Engineering Manifesto

Observability is part of done

Structured logs, metrics, and traces ship with the feature — not as a follow-up ticket after the outage.

Principle

If you cannot see how a change behaves in production, you did not finish it. Observability is a definition-of-done item for service work.

Therefore we practice

  • Emit structured logs (fields, not only prose) with correlation/causation ids across hops
  • Expose golden signals for the journey: rate, errors, latency (and saturation where relevant)
  • Propagate trace context through sync and async boundaries
  • Log outcomes at decision points (authorize, capture, route) — not only stack traces
  • Dashboards and alerts exist before or with the rollout, not “later”

Smells

  • log.info("error " + e) with no ids, no structure, no context
  • Only CPU/memory charts when users report failed checkouts
  • New endpoints with zero metrics
  • Debugging by SSH into a random replica

← Software Engineering Manifesto