Reliability and operability practices
Designing for failure — runbooks, deployments, capacity, and production readiness as engineering work.
DevOps reliability is engineering, not hope. Operability is designed in before the first page.
Production readiness
Before calling a service “done”: health/readiness, metrics and logs, alerts with runbooks, rollback path, dependency timeouts, and capacity assumptions written down. A feature flag without an owner is not readiness.
Safe change
Prefer progressive delivery, migrations that expand/contract, and feature flags with cleanup. Match change size to error budget and blast radius.
Incident practice
Declare early, communicate clearly, mitigate first, then diagnose. Templates beat improvisation. Game days and failure injection find gaps before customers do.
Capacity and cost
Watch saturation (CPU, threads, connections, lag, disk). Autoscale with ceilings. Cost is an operability signal — surprise bills often mean runaway retries or forgotten environments.
Toil reduction
Automate repetitive manual ops. If humans paste digests into three consoles every release, that is a backlog item — not a badge of honor.