Shards, replicas, and capacity
How primary shards, replicas, and heap sizing decide latency, resilience, and whether rebalancing helps.
Cluster topology is a capacity plan expressed as JSON. Wrong shard counts haunt you for years.
Primaries set the ceiling
Primary shard count is chosen at index creation (for practical purposes). Too few and you cannot parallelize. Too many and you pay cluster-state and per-shard overhead for tiny shards.
Rule of thumb direction: aim for shard sizes in a healthy band for your version and heap — not “one shard per day forever” without ILM, and not thousands of tiny shards.
Replicas buy reads and availability
Replicas take search load and survive node loss. They cost disk and indexing CPU. number_of_replicas: 0 on production search is a choice you make with eyes open.
Hot spots
Uneven document routing or aggressive custom routing creates hot primary shards. More nodes will not fix a single hot key. Fix the key or split the index design.
Heap and circuit breakers
Fielddata, aggs, and bulk indexing compete for heap. Stay on documented heap guidance; watch circuit-breaker trips as product bugs, not noise.
Plan reindex and rollover
Capacity includes how you grow: aliases, rollover (ILM), reindex windows, and snapshot restore drills. A cluster that cannot restore is not durable — it is hopeful.