Operating pipelines
Speed, ownership, templates, and metrics — keeping CI/CD useful after the first green build.
Pipelines rot like any other system. Day-2 CI/CD is about feedback speed and clear ownership.
Measure what hurts
Track:
- time from commit to green / to production
- flake rate and rerun rate
- queue time on runners
- change failure rate and time to restore
If PR feedback takes 40 minutes, people batch work and skip runs. Speed is a reliability feature.
Templates without captivity
Platform provides paved-road workflows; services may need escape hatches. A single mega-pipeline forced on every repo creates tickets forever. Version shared templates and changelog them.
Own the red build
Pager or Slack ownership for broken main. “Someone will notice” is not a strategy.
Change the pipeline deliberately
Pipeline edits go through the same review path. Shadow or canary new shared workflows before forcing them org-wide.
Keep humans for judgment, not toil
Automate deploys, evidence collection, and rollbacks. Reserve humans for risk decisions and incident response — not copying digests into three consoles.
Healthy CI/CD feels boring: predictable duration, rare flakes, obvious promotion, and a known owner when it breaks.