Learning · SLI/SLO

Alerting on burn

Multi-window burn-rate alerts — paging on meaningful SLO risk, not on every blip.

Good SLO alerting asks: Are we burning error budget fast enough that we will miss the objective without action?

Burn rate over raw thresholds

A short spike may be fine inside the monthly budget. A sustained elevated error rate may exhaust the budget before anyone notices a flat “error > 1%” page that people muted.

Multi-window burn-rate alerts (fast window + slow window) catch both “melting now” and “dying by a thousand cuts.”

Page humans for action

Alerts should imply a next step: mitigate, roll back, shed load, declare incident. Tickets for slow burns; pages for fast burns that threaten the SLO soon.

Symptom first, cause second

Page on SLI burn. Use CPU, saturation, and dependency errors as diagnostic dashboards once humans are awake. Cause-based pages without symptom context create alert fatigue.

Tune with retrospectives

Every false page and every missed burn belongs in a short review: wrong SLI, wrong window, wrong threshold, or missing dependency. Alert quality is an engineering backlog item.

Synthetic probes

Synthetics catch total outages when traffic is low. They complement — they do not replace — user-traffic SLIs. Keep them honest about geography and auth paths.

← SLI/SLO