Learning · SLI/SLO

SLI/SLO

Reliability needs a contract with users, not a wall of green graphs. These notes cover how to define SLIs, set SLOs, spend error budgets, and alert on burn — so on-call wakes up for user pain, not CPU vibes.

Topics

  • What SLIs and SLOs are for

    The reliability contract — indicators, objectives, and why “uptime 99.9%” alone is not a strategy.

  • Choosing good SLIs

    Availability, latency, quality, and freshness — picking indicators that track real user journeys.

  • Setting SLO targets

    How ambitious to be — windows, percentiles, and targets grounded in history and product risk.

  • Error budgets

    How much unreliability you can spend — using budget to balance velocity and reliability.

  • Alerting on burn

    Multi-window burn-rate alerts — paging on meaningful SLO risk, not on every blip.

  • Dependencies and multi-service SLOs

    How to think about SLOs when your service calls others — and when platform SLIs matter.

  • Reporting and culture

    Making SLOs visible — reviews, release freezes, and reliability as a shared product concern.

  • Anti-patterns

    Dashboard theater, unowned SLOs, and other ways reliability metrics lose their meaning.