SLI/SLO
Reliability needs a contract with users, not a wall of green graphs. These notes cover how to define SLIs, set SLOs, spend error budgets, and alert on burn — so on-call wakes up for user pain, not CPU vibes.
Topics
- What SLIs and SLOs are for
The reliability contract — indicators, objectives, and why “uptime 99.9%” alone is not a strategy.
- Choosing good SLIs
Availability, latency, quality, and freshness — picking indicators that track real user journeys.
- Setting SLO targets
How ambitious to be — windows, percentiles, and targets grounded in history and product risk.
- Error budgets
How much unreliability you can spend — using budget to balance velocity and reliability.
- Alerting on burn
Multi-window burn-rate alerts — paging on meaningful SLO risk, not on every blip.
- Dependencies and multi-service SLOs
How to think about SLOs when your service calls others — and when platform SLIs matter.
- Reporting and culture
Making SLOs visible — reviews, release freezes, and reliability as a shared product concern.
- Anti-patterns
Dashboard theater, unowned SLOs, and other ways reliability metrics lose their meaning.