Learning · SLI/SLO

Choosing good SLIs

Availability, latency, quality, and freshness — picking indicators that track real user journeys.

A bad SLI is worse than none: it creates confident wrong decisions. Start from user journeys, not from whatever Prometheus already scrapes.

Common SLI types

  • Availability / success rate — requests that succeed by the user’s definition (not merely HTTP 200 from a health check)
  • Latency — how long a meaningful operation takes (p95/p99 of checkout, search, capture)
  • Quality / correctness — results that meet expectations (search relevance proxies, reconciliation match rate)
  • Freshness / lag — how stale a projection, index, or consumer lag is relative to the source of truth

Pick a few that matter. Ten SLIs dilute ownership.

Measure at the edge of the promise

Prefer indicators close to the user or the API they call. Internal queue depth can explain pain; it rarely is the SLI unless the product promise is “messages processed within N minutes.”

For async systems, define success as “fact applied within SLO,” not “message accepted by Kafka.”

Exclude what misleads

Client cancels, bot traffic, and synthetic probes may need separate series or filters. Mixing them with real user traffic turns “availability” into fiction.

Document the formula

Write the numerator, denominator, window, and exclusions in one place. “Success rate” without a formula is a hallway argument waiting to happen.

← SLI/SLO