Learning · SLI/SLO

Anti-patterns

Dashboard theater, unowned SLOs, and other ways reliability metrics lose their meaning.

Common ways SLI/SLO programs rot — and what to do instead.

SLOs on vanity metrics

CPU, pod count, or “instances up” as the primary SLO. Instead: user-journey success and latency.

Too many SLOs

Everything is 99.99% and nothing is owned. Instead: a handful of journey SLOs with names on them.

Alerting on causes only

Page on disk, GC, and replica lag with no symptom link. Instead: burn on SLIs; causes in playbooks.

Moving the goalposts after outages

Quietly exclude the bad week. Instead: spend the budget, then improve the system or renegotiate the target in the open.

SLOs without consequences

Budget hits zero and releases continue unchanged. Instead: a written policy that actually changes process.

Copy-pasted 99.9%

Same target on a batch job and a checkout API. Instead: targets matched to stakes and historical reality.

Health-check availability

Load balancer ping success as “uptime.” Instead: measure the real API or journey, including dependency failures users feel.

If your SLO program cannot change a release decision, it is decoration.

← SLI/SLO