Anti-patterns
Dashboard theater, unowned SLOs, and other ways reliability metrics lose their meaning.
Common ways SLI/SLO programs rot — and what to do instead.
SLOs on vanity metrics
CPU, pod count, or “instances up” as the primary SLO. Instead: user-journey success and latency.
Too many SLOs
Everything is 99.99% and nothing is owned. Instead: a handful of journey SLOs with names on them.
Alerting on causes only
Page on disk, GC, and replica lag with no symptom link. Instead: burn on SLIs; causes in playbooks.
Moving the goalposts after outages
Quietly exclude the bad week. Instead: spend the budget, then improve the system or renegotiate the target in the open.
SLOs without consequences
Budget hits zero and releases continue unchanged. Instead: a written policy that actually changes process.
Copy-pasted 99.9%
Same target on a batch job and a checkout API. Instead: targets matched to stakes and historical reality.
Health-check availability
Load balancer ping success as “uptime.” Instead: measure the real API or journey, including dependency failures users feel.
If your SLO program cannot change a release decision, it is decoration.