Learning · Elasticsearch

Querying and relevance

Filters vs queries, ranking, and keeping search results useful instead of merely matching.

Matching documents is easy. Returning the right documents in the right order is the product.

Filter first, score second

Use filters for hard constraints (status, tenant, range). They cache well and do not need scores. Use queries when relevance ranking matters.

A common pattern: bool with filter clauses for security and facets, and must / should for text relevance.

Know your query types

  • match / multi_match for analyzed text
  • term / terms for exact keyword values
  • range for dates and numbers
  • match_phrase when order matters
  • function score / rank features when business boosts matter

Picking match when you meant term is a frequent source of “why is this id not found?”

Relevance is a product loop

Defaults are a starting point. Tune with:

  • real search logs and click data
  • synonyms and field boosts grounded in user language
  • separate ranking for autocomplete vs full search
  • protected tests so a mapping change does not tank precision

Do not ship unexplained script scores that only one engineer understands.

Pagination and deep pages

from + size gets expensive deep in the result set. Prefer search_after (or point-in-time) for deep pagination and exports. Cap what the UI offers; nobody needs page 500 of junk.

Security in the query

Every search path must enforce tenant / authz filters. A missing filter is a data leak. Build that into a shared client or search template — not each caller’s goodwill.

← Elasticsearch