Learning · Elasticsearch

Mappings and analysis

How field types and analyzers decide recall, precision, and whether your index can evolve safely.

The mapping is the schema. Get it wrong and you fight relevance, aggregations, and storage forever — or you reindex under pressure.

Explicit beats dynamic

Dynamic mapping is fine for spikes. Production indexes need explicit mappings for fields you query, sort, or aggregate on.

Decide per field:

  • text vs keyword (full-text vs exact / aggregations)
  • numeric / date / boolean for range and sort
  • whether you need both (text + keyword multi-fields is common)

Surprises like a numeric id mapped as text break sorting and aggregations quietly.

Analyzers tokenize and normalize text at index and query time. Choices matter:

  • language stemmers vs exact tokens
  • lowercase, ascii folding, synonyms
  • edge n-grams for autocomplete (costly if overused)

Index-time and query-time analysis must match the product intent. Searching with a different analyzer than you indexed with is a classic “works in Kibana, fails in the app” bug.

Nested and object data

Flattened objects are simple until you need independent matches inside arrays. nested types fix false matches across objects — at a cost in memory and query complexity. Use them when correctness demands it, not by default.

Evolving mappings

Many mapping changes are not in-place compatible. Plan for:

  • additive new fields when possible
  • reindex into a new index + alias swap for breaking changes
  • index templates so new indices inherit the right schema

A mapping review belongs in design review — next to API contracts — not as a Friday hotfix.

← Elasticsearch