Learning · SLI/SLO

What SLIs and SLOs are for

The reliability contract — indicators, objectives, and why “uptime 99.9%” alone is not a strategy.

SLIs measure what users experience. SLOs are the targets you commit to for those measures. Together they turn reliability from opinion into engineering.

The stack of meaning

  • SLI — a quantitative measure of one aspect of service (e.g. fraction of successful payment captures, p99 latency of search)
  • SLO — the target for that SLI over a window (e.g. 99.9% success over 28 days)
  • SLA — an external/business commitment, often with penalties; usually looser or different from internal SLOs
  • Error budget — how much unreliability you can “spend” before the SLO is missed

Seniors design SLIs/SLOs for product decisions (ship vs freeze, page vs ticket). SLAs are legal/commercial; do not confuse the two.

What they replace

“The dashboard looks fine,” “CPU is low,” and “we had no Sev-1s this week” are weak substitutes. Users care whether their request worked, fast enough, with fresh enough data.

What they do not replace

SLOs do not invent architecture, fix bad boundaries, or excuse missing runbooks. They focus attention and tradeoffs. A perfect SLO on a wrong SLI still pages the wrong people.

Smell test

Ask: If this number is red, would a thoughtful user agree something is wrong? If no, redesign the indicator before you automate alerts.

← SLI/SLO