Choosing good SLIs
Availability, latency, quality, and freshness — picking indicators that track real user journeys.
A bad SLI is worse than none: it creates confident wrong decisions. Start from user journeys, not from whatever Prometheus already scrapes.
Common SLI types
- Availability / success rate — requests that succeed by the user’s definition (not merely HTTP 200 from a health check)
- Latency — how long a meaningful operation takes (p95/p99 of checkout, search, capture)
- Quality / correctness — results that meet expectations (search relevance proxies, reconciliation match rate)
- Freshness / lag — how stale a projection, index, or consumer lag is relative to the source of truth
Pick a few that matter. Ten SLIs dilute ownership.
Measure at the edge of the promise
Prefer indicators close to the user or the API they call. Internal queue depth can explain pain; it rarely is the SLI unless the product promise is “messages processed within N minutes.”
For async systems, define success as “fact applied within SLO,” not “message accepted by Kafka.”
Exclude what misleads
Client cancels, bot traffic, and synthetic probes may need separate series or filters. Mixing them with real user traffic turns “availability” into fiction.
Document the formula
Write the numerator, denominator, window, and exclusions in one place. “Success rate” without a formula is a hallway argument waiting to happen.