Indexing and ingestion
Bulk indexing, ids, routing, and keeping the search index in sync with the system of record.
Search quality starts with how documents arrive. Slow or inconsistent ingestion shows up as stale results and angry users.
Prefer bulk, not chatty singles
Use the bulk API for steady pipelines. Tune bulk size to document size and heap — too large stalls; too small wastes overhead. Watch indexing rejections and bulk reject metrics.
Stable document ids
Choose ids you can update and delete later (business keys or deterministic hashes). Random ids force awkward deletes and duplicates when the same entity is indexed twice.
Routing and parent locality
Custom routing can keep related documents on the same shard for certain join patterns — and create hot shards if the key is skewed. Default routing by id is safer until you have a measured reason to customize.
Near-real-time vs durable truth
A successful index response does not mean every replica searcher sees the doc instantly. Understand refresh (search visibility) vs flush/translog (durability). Product copy that says “instant” needs an honest latency budget.
Keep the pipeline rebuildable
Whether you use Logstash, Beats, Kafka Connect, or a custom indexer:
- make runs idempotent (same doc id → upsert)
- checkpoint progress
- support full reindex from the source of truth
- separate hot write aliases from long reindex jobs when possible
Deletes and updates
Updates are often retrieve-merge-reindex under the hood. High-churn fields may belong in a different modeling approach (or fewer partial updates). Soft deletes in the source must map to real deletes or filters in ES — or tombstones pile up until merge pressure hurts.
Ingestion is successful when the index is correct, timely enough, and recoverable — not merely when the bulk call returned 200.