Learning · Software Engineering Manifesto

Make failure explicit

Timeouts, budgets, and degradation are design — infinite waits and blind retries are not resilience.

Principle

Every remote dependency can be slow or down. Failure must be a designed state with bounds, not an accident discovered by exhausted thread pools.

Therefore we practice

  • Set timeouts from dependency behavior and user SLOs — never “default forever”
  • Bound retries with backoff and jitter; cap attempts; distinguish retryable vs not
  • Use bulkheads/limits so one dependency cannot consume all capacity
  • Define degraded behavior when a non-critical dependency fails
  • Prefer fail fast on the critical path over hanging the user

Smells

  • Retry storms that amplify an outage
  • Liveness probes that call dependencies and restart healthy pods
  • No timeout on “important” calls because they “must succeed”
  • Cascading failure across a synchronous chain of five services

← Software Engineering Manifesto