Operating Elasticsearch
ILM, monitoring, upgrades, and the signals that mean search is about to miss its SLO.
Day-2 Elasticsearch is mostly discipline: keep indices manageable, watch the right metrics, and practice recovery.
Index lifecycle
Use ILM (or equivalent) for time-series and growing datasets:
- hot for fresh writes and low-latency search
- warm/cold for cheaper storage when latency can relax
- delete when retention expires
Unbounded indices without rollover are a slow-moving outage.
Watch what users feel
Alert on:
- search latency (p95/p99) and error rates
- indexing latency and bulk rejections
- JVM heap pressure and circuit breakers
- disk watermarks and relocation storms
- pending tasks / cluster status yellow-red lasting beyond expected recovery
Yellow after a node restart can be normal briefly. Yellow forever is not.
Snapshots are mandatory
Schedule snapshots to durable storage. Test restore into a scratch cluster. An untested snapshot is fiction.
Upgrades and compatibility
Rolling upgrades need version compatibility awareness. Reindex-for-upgrade paths exist for a reason — budget them. Keep clients on supported protocol versions.
Secure the HTTP/API surface
TLS, authn/authz, least-privilege API keys, and network isolation. Open Elasticsearch on the public internet is still an embarrassing class of incident.
Operate so rebuild, restore, and rollover are boring — that is when search stays trustworthy.