Learning · Software Engineering Manifesto
Measure what users feel
KPIs and SLIs track user journeys — vanity infrastructure metrics do not replace a reliability contract.
Principle
We manage what we measure. Prefer indicators tied to user success and experience (and explicit SLOs/KPIs) over green dashboards that ignore pain.
Therefore we practice
- Define measurable outcomes for critical journeys (success rate, latency, freshness)
- Agree SLO/KPI targets and error budgets where reliability matters
- Alert on burn against those targets, not on every blip of CPU
- Review metrics in planning and incidents — numbers that never change decisions are decoration
- Separate product KPIs from pure infra gauges; use both, confuse neither
Smells
- “Uptime 100%” while the checkout API fails open for users
- No latency objective on a user-facing API
- Alerts everyone ignores
- Shipping features while error budget is exhausted with no policy change