Learning · Software Engineering Manifesto
Make failure explicit
Timeouts, budgets, and degradation are design — infinite waits and blind retries are not resilience.
Principle
Every remote dependency can be slow or down. Failure must be a designed state with bounds, not an accident discovered by exhausted thread pools.
Therefore we practice
- Set timeouts from dependency behavior and user SLOs — never “default forever”
- Bound retries with backoff and jitter; cap attempts; distinguish retryable vs not
- Use bulkheads/limits so one dependency cannot consume all capacity
- Define degraded behavior when a non-critical dependency fails
- Prefer fail fast on the critical path over hanging the user
Smells
- Retry storms that amplify an outage
- Liveness probes that call dependencies and restart healthy pods
- No timeout on “important” calls because they “must succeed”
- Cascading failure across a synchronous chain of five services