Learning · Elasticsearch

Operating Elasticsearch

ILM, monitoring, upgrades, and the signals that mean search is about to miss its SLO.

Day-2 Elasticsearch is mostly discipline: keep indices manageable, watch the right metrics, and practice recovery.

Index lifecycle

Use ILM (or equivalent) for time-series and growing datasets:

  • hot for fresh writes and low-latency search
  • warm/cold for cheaper storage when latency can relax
  • delete when retention expires

Unbounded indices without rollover are a slow-moving outage.

Watch what users feel

Alert on:

  • search latency (p95/p99) and error rates
  • indexing latency and bulk rejections
  • JVM heap pressure and circuit breakers
  • disk watermarks and relocation storms
  • pending tasks / cluster status yellow-red lasting beyond expected recovery

Yellow after a node restart can be normal briefly. Yellow forever is not.

Snapshots are mandatory

Schedule snapshots to durable storage. Test restore into a scratch cluster. An untested snapshot is fiction.

Upgrades and compatibility

Rolling upgrades need version compatibility awareness. Reindex-for-upgrade paths exist for a reason — budget them. Keep clients on supported protocol versions.

Secure the HTTP/API surface

TLS, authn/authz, least-privilege API keys, and network isolation. Open Elasticsearch on the public internet is still an embarrassing class of incident.

Operate so rebuild, restore, and rollover are boring — that is when search stays trustworthy.

← Elasticsearch